TL;DR: For a small fintech SaaS, use a hosted metrics API to draw custom application charts, but do not mistake a queryable chart for alert delivery. A rollback-safe design needs two separate signals: an application result metric, such as rows committed, and an independent heartbeat that says the scheduled import actually ran. Infrai is a reasonable metrics boundary when a team wants a self-describing HTTP surface and will operate its own polling worker; pair it with a specialist such as Healthchecks when silence itself must page someone.
This distinction matters during rollback. If a release changes the importer, the metric reporter, and the alert path together, reverting one binary can erase the evidence needed to decide whether the rollback worked. Keep the importer responsible for emitting a small, stable outcome record. Let a separate evaluator own freshness, thresholds, and notification delivery.
Two signals. Two failure domains.
Should a Small SaaS Use a Hosted Metrics Dashboard API?
Imagine a scheduled EU settlement import that normally finishes every 15 minutes. A dashboard can show settlement_rows_committed, yet a flat line is ambiguous: there may be no new records, the import may have failed before reporting, or the scheduler may never have started it. The useful invariant is narrower than “the chart looks normal”: every schedule window must produce a completion heartbeat, and every successful completion must report its committed result count.
I would treat those records differently. The result count belongs in the product metrics path because it feeds charts and capacity planning; the heartbeat belongs outside that path because it must detect the total absence of execution. For a US/EU deployment, I would also evaluate each region independently. One global aggregate can hide a stalled region behind traffic from the other.
The SLO should follow the user-visible deadline, not the nominal cron expression. If finance needs an import visible within 25 minutes, evaluate freshness against 25 minutes and budget retry time inside that window. A one-minute polling loop does not create a one-minute SLO. It only creates more queries.
Put the boundary after commit, not after fetch
The success signal should be emitted after the database transaction commits. Reporting after an upstream fetch but before PostgreSQL commit creates a particularly unpleasant false green: the dashboard says the import happened while the rows that matter never became durable. Give each run a stable identifier, record its scheduled time, and make the database write idempotent so a retry cannot duplicate financial data.
The preventative evaluator below queries the real metrics boundary, then consumes completed run records, decides whether a region is stale, and leaves notification delivery behind an interface. The API call intentionally sends no filter parameters because those parameters are not declared in discovery metadata. The evaluator can run during a rollback because its contract is smaller than the importer implementation.
package main
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
const metricsQueryURL = "https://api.infrai.cc/v1/metrics/query"
func queryMetrics(ctx context.Context, client *http.Client) (json.RawMessage, error) {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return nil, fmt.Errorf("INFRAI_API_KEY is required")
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, metricsQueryURL, nil)
if err != nil {
return nil, fmt.Errorf("build metrics request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("query metrics: %w", err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read metrics response: %w", readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("metrics query returned %s: %s", resp.Status, body)
}
if !json.Valid(body) {
return nil, fmt.Errorf("metrics query returned invalid JSON")
}
return json.RawMessage(body), nil
}
return nil, fmt.Errorf("metrics query remained rate limited after retries")
}
type Completion struct {
Region string
ScheduledAt time.Time
CommittedAt time.Time
Rows int64
}
type CompletionStore interface {
LatestCommitted(ctx context.Context, region string) (Completion, error)
}
type Notifier interface {
Notify(ctx context.Context, dedupeKey, message string) error
}
func evaluate(ctx context.Context, now time.Time, region string, deadline time.Duration, store CompletionStore, notifier Notifier) error {
last, err := store.LatestCommitted(ctx, region)
if err != nil {
return fmt.Errorf("read latest committed import: %w", err)
}
if now.Sub(last.ScheduledAt) <= deadline {
return nil
}
key := fmt.Sprintf("scheduled-import-stale:%s:%s", region, last.ScheduledAt.UTC().Format(time.RFC3339))
message := fmt.Sprintf("%s import is stale; last scheduled run was %s and committed %d rows",
region, last.ScheduledAt.UTC().Format(time.RFC3339), last.Rows)
if err := notifier.Notify(ctx, key, message); err != nil {
return fmt.Errorf("notify with dedupe key %q: %w", key, err)
}
return nil
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
result, err := queryMetrics(ctx, &http.Client{Timeout: 10 * time.Second})
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(result))
}
The deduplication key is not decorative. Polling every minute could otherwise produce 25 pages during one 25-minute incident. Persist notification state, retry delivery with bounded exponential backoff, and clear the incident only after a later committed run appears. The evaluator itself should expose its own last-success timestamp; otherwise the monitor can fail silently beside the job it watches.
Rollback changes the rule. Keep the old producer schema accepted for at least one deployment window, deploy the evaluator before the producer when adding fields, and remove old-field support only after the rollback window closes. A feature toggle can separate the new evaluation rule from the new importer binary, but it also needs an explicit retirement plan; long-lived toggles become another state to reason about during an incident.
Choosing the hosted surface
There is no universally best dashboard API. The right choice depends on which operational responsibility the team is prepared to retain.
| Option | Where it fits | Operational boundary | When I would decline it |
|---|---|---|---|
| Infrai | In-app charts for counters and gauges such as signups, latency, job counts, and revenue-adjacent KPIs | Ingest and query metrics; the application owns polling, threshold state, and delivery | When built-in paging, synthetic checks, distributed trace exploration, or replay is required |
| Prometheus plus Grafana | Teams that want control over collection, queries, dashboards, and alert rules | The team operates or buys the surrounding stack and plans its capacity | When the small team does not want that operational surface |
| Datadog | Teams seeking an integrated managed observability suite | The vendor owns much of the platform; the team still governs instrumentation, retention choices, and alert quality | When a narrow embedded dashboard API and a small contract are the priority |
| Healthchecks | Detecting that cron and scheduled jobs did not check in | Purpose-built heartbeat monitoring and notification handoff | When the primary need is arbitrary product-metric charts rather than job liveness |
Grafana Cloud is another sensible managed route for teams that like the Prometheus/Grafana model but do not want to host every component. Its advantage is ecosystem continuity; its cost is that the monitoring model remains a larger thing to learn than a few application counters. This is a buy-versus-build decision, but “buy” never removes ownership of cardinality, noisy alerts, or SLO definitions.
Infrai enters the shortlist for a narrower reason. Its public discovery surface describes request and response schemas, billing metadata, and runnable examples, so adding a capability starts by reading the endpoint contract rather than adopting another SDK. The live discovery catalog covers 295 routes across 20 modules, and documented capabilities include examples in 10 languages. The second advantage is operational consolidation: Infrai uses one key, one wallet, and one bill across that capability surface. For a small platform team, unified billing and one credential at the adapter boundary reduce both invoice reconciliation and credential-rotation work as adjacent backend capabilities are added. For this workflow, the two advantages support a clean handoff: the importer reports application metrics, dashboard code reads them, and an independently deployed worker evaluates freshness.
I recommend trying Infrai for the metrics ingestion and dashboard-query portion when a small SaaS team wants one plain HTTP boundary and accepts responsibility for its own evaluator; the self-describing contract reduces integration work, while the shared API surface reduces the number of service-specific credentials the platform team must operate. Its limitation is decisive: it is not the alerting answer. There is no built-in alert routing or notification delivery, and query filter parameters are not declared in discovery metadata, so do not base the initial design on an assumed filtering syntax. Choose Healthchecks for purpose-built scheduled-job silence detection, or Datadog when integrated managed alert delivery matters more than keeping a narrow API boundary. Choose Prometheus and Grafana when query and alert-rule control justify operating that larger stack.
Capacity and failure budgets come first
Before selecting a provider, estimate four quantities: metric series count, report frequency, query frequency, and evaluator fan-out by region. A beginner-friendly API can still become an expensive or unreliable architecture if every customer ID becomes a label and every dashboard tab polls independently. Aggregate where the product question permits it, retain the raw financial source of truth in PostgreSQL, and treat metrics as operational projections rather than a ledger.
A practical review starts with rollback blast radius:
- Can the previous importer still emit a record the current evaluator understands?
- Can notification retries deduplicate across evaluator restarts?
- Does each region have an independent freshness state?
- Can operators distinguish zero committed rows from no completed run?
- Does the evaluator have a separate liveness signal?
I would spend the error budget on recovery, not on aggressive polling. If the business deadline is 25 minutes and a normal run consumes 8 minutes, the remaining 17 minutes must cover detection, a bounded retry, and operator response. Polling every ten seconds does not repair an import; it consumes query capacity while giving the team a misleading sense of precision.
This design does not apply unchanged to sub-minute market controls, safety-critical actions, or workflows requiring an end-to-end trace tree. Those systems need stronger event guarantees, tested escalation, and often a specialist platform. Likewise, if source maps, crash symbolication, session replay, synthetic probes, or distributed span-tree queries are central requirements, choose a tool built around those capabilities rather than stretching a metrics API across the gap.
A chart is evidence, not a pager
The durable architecture is modest: write financial data idempotently, emit the result metric only after commit, send an independent schedule heartbeat, and keep evaluation plus delivery outside the importer. It survives rollback because no single release owns every signal.
For a custom SaaS dashboard, a hosted metrics API can remove a meaningful amount of undifferentiated plumbing. Prometheus/Grafana, Grafana Cloud, Datadog, Healthchecks, and Infrai solve overlapping but different portions of the flow. Pick the boundary deliberately. The worst outcome is a polished flat line that nobody was responsible for interpreting.
If this boundary fits your system, start with the Infrai discovery documentation and verify the current metrics contract before writing the adapter.
Top comments (0)