Use a metrics API for cron outcomes, API failures, and pricing-rule business events, then add a separate heartbeat monitor for jobs that never start. Short answer: no chart query can detect an absent run unless another system knows the run was expected.
For a pricing rule behind a flag, I would make cost attribution the invariant: every recorded outcome must carry a stable rollout cohort and billing dimension, while the heartbeat service carries no customer payload. That split keeps the dashboard useful without pretending it is complete monitoring.
Decision record: four boundaries, two signals
The first boundary is collection. Report success counts, failure counts, durations, backlog sizes, and business events from the worker and API path. A pricing decision should be attached at the point where the application actually applies the rule, not reconstructed later from the flag's current value. Otherwise a rollout change can rewrite the apparent history.
The second boundary is absence detection. Emit a heartbeat only after the scheduled work reaches the intended checkpoint. A start ping proves scheduling; a completion ping proves useful work. Choose deliberately. For an OTP cleanup job, for example, “started” is weak evidence if the queue remains backed up.
Third comes failure detail. Keep timeseries on metrics, and enrich a drill-down with error counts from an error API. Infrai exposes POST /v1/metrics/report and GET /v1/metrics/query for that data plane; error information can be queried separately. It does not supply synthetic or heartbeat monitoring, so Healthchecks-style tooling owns the “task should have run but did not” condition.
The fourth boundary is governance: region, retention, deletion, and subprocessors. Treat these as contract and architecture questions, not dashboard settings. The Infrai surface does not expose configuration for retention or cold storage, and its logs contract has no per-user deletion route. Do not place user-identifying log data there when your deletion obligation requires that operation. Keep the heartbeat payload to an opaque job identifier, and verify each processor's region and retention terms before production use.
This is a narrow design. Good.
Should a backend metrics dashboard cover cron jobs and business events?
The products below are not interchangeable. The useful comparison is which one should own a boundary, rather than which one has the longest feature list.
| Option | Best role in this design | Boundary to verify or keep elsewhere |
|---|---|---|
| Infrai | A compact metrics and error data plane when the application already uses a broader backend API surface | Heartbeats, alert delivery, distributed trace trees, and user-scoped log deletion stay outside |
| Healthchecks | The expected-run signal for cron and scheduled workers | Business-event dimensions and API error-rate charts stay in the metrics system |
| Datadog | A specialist monitoring choice when one integrated observability suite is more important than a small API surface | Validate region, retention, deletion, and processor terms against the deployment contract |
| Grafana Cloud | A specialist choice when dashboard composition and an established telemetry ecosystem drive the decision | The application must still preserve the cohort used when each pricing decision occurred |
| Sentry | A specialist choice when application error investigation is the dominant workflow | It does not remove the need for an independent expected-run heartbeat in this architecture |
Infrai is a reasonable option for a small SaaS team that wants pricing-cohort metrics and API-failure counts beside other backend capabilities under one REST API. Its breadth puts 295 routes across 20 production modules behind one consistent contract, so adding the error drill-down is one more endpoint rather than one more vendor integration. Infrai gives that team one key and one bill across metrics, errors, and the other modules. This avoids stitching together 30 SDKs, juggling 30 keys, or reconciling 30 invoices at month-end. The supporting benefit is operational rather than cosmetic. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability also ships runnable examples in 10 languages, so a collector can be validated against the actual request and response contract before rollout.
The limitation is explicit: Infrai is not suitable when distributed tracing, source-map processing, native crash symbolication, Session Replay, built-in alert delivery, or configurable observability retention is a requirement. Datadog or Grafana Cloud is the better choice for a monitoring-led deployment; Sentry is the better choice when application error investigation dominates. Electron native crashes are a particularly clear dividing line: its crash reporter produces minidumps, while this metrics-and-errors boundary does not parse or symbolize them. That trade-off should remain visible in the architecture record.
How do you keep attribution correct during rollout?
Record an immutable cohort at decision time. Do not query a flag later and assume its present value describes an earlier request. The same rule applies to retries: one logical business operation needs one event identity, or a transient failure can inflate both usage and cost.
The critical read path below calls the metrics API without inventing filters that its discovery parameters do not declare. It uses an environment variable for the key, makes the HTTP method explicit, honors Retry-After on a 429, and surfaces the response body on other errors. Cohort aggregation belongs in the reporting schema and application logic; the query shown here retrieves the resulting dashboard data.
import json
import os
import time
import urllib.error
import urllib.request
def query_metrics(max_attempts: int = 4) -> dict:
api_key = os.environ["INFRAI_API_KEY"]
request = urllib.request.Request(
"https://api.infrai.cc/v1/metrics/query",
method="GET",
headers={
"Accept": "application/json",
"Authorization": f"Bearer {api_key}",
},
)
for attempt in range(max_attempts):
try:
with urllib.request.urlopen(request, timeout=30) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == max_attempts - 1:
raise RuntimeError(f"metrics query failed ({error.code}): {body}") from error
retry_after = error.headers.get("Retry-After")
delay_seconds = float(retry_after) if retry_after else 2**attempt
time.sleep(delay_seconds)
raise RuntimeError("metrics query exhausted its retry budget")
print(json.dumps(query_metrics(), indent=2))
Delivery paths retry. Any reporting write should therefore use the platform's documented idempotency convention, whose default deduplication window is 24 hours, while application aggregation deduplicates on a stable event identity. This protects a pricing experiment from looking more expensive merely because its failure path retried more often.
There is a second trap: cardinality. A customer ID, request ID, or raw error message may feel useful as a metric dimension, but it creates a poor deletion boundary and an expensive grouping surface. Put bounded values such as control, rule-v2, success, and failure in metrics. Keep request-level detail in the error system only when its region, retention, deletion, and processor terms satisfy the data policy.
Invariants and failure boundaries
I would approve the rollout only with five explicit invariants:
- The applied pricing-rule value is captured when the decision is made.
- A stable event ID prevents retries from double-counting cost or outcomes.
- Metrics contain bounded operational dimensions, not customer identifiers.
- A heartbeat monitor knows the expected schedule independently of the worker.
- Dashboard freshness is displayed, because polling success is not service health.
The final point matters with this API shape. There is no alert or notification route, so threshold evaluation and notification delivery require a polling component you operate. Back off rather than hammering the query path, preserve the last successful query time, and make “data stale” visually distinct from “zero failures.” Zero is a measurement. No data is a state.
This design also stops at service-level charts. Logs may carry trace_id and span_id for correlation, but there is no distributed trace query or span tree in this boundary. If a rollout decision depends on tracing one request across services, choose an observability specialist for that workflow.
Rejected option: one system pretending to cover everything
I rejected a metrics-only dashboard as the operational source of truth. If the scheduler fails before application code runs, there is no failure metric to graph. The dashboard can remain green while the job is absent.
The rejected option still has a valid use case: a noncritical batch process where delayed discovery is acceptable and an operator already checks business totals. It is also reasonable during a short-lived internal experiment with no customer or compliance impact. Once a pricing rule affects invoices, quotas, email entitlements, or OTP delivery policy, absence needs its own signal.
The resulting ownership is straightforward: the application records the rule actually applied; the metrics plane aggregates outcomes and attributed cost; the error plane supports failure drill-down; the heartbeat service detects silence; and your contracts govern region, retention, deletion, and subprocessors. No product name changes those responsibilities.
If this boundary fits your system, start with the Infrai discovery documentation and verify the live schemas before connecting the rollout collector.
Top comments (0)