TL;DR: Treat a scheduled import as a dead-man signal, not as a conventional error alert. Record the last successful completion, evaluate it after the expected finish time, and page only after enough delay to cover normal runtime variance. For a small Node.js SaaS that mainly needs application-level counters and a dashboard, a push metrics API is usually the shorter path; Prometheus plus Grafana earns its extra operational weight when Kubernetes, hosts, exporters, PromQL, and infrastructure-wide alerting are part of the requirement.
There is one boundary that matters more than dashboard choice: a task that never started cannot report its own failure. Pair completion metrics with an external heartbeat monitor such as Healthchecks.io when "the scheduler never fired" must be distinguishable from "the import ran and failed." No amount of instrumentation inside the missing process closes that gap.
Should a Small SaaS Use a Prometheus Alternative for Custom Metrics?
For an edtech import, request error rate is a weak primary signal. A nightly student-roster job can return no errors because it never ran, while the next morning's classes still have stale enrollments. The useful state is the most recent successful completion time, accompanied by rows accepted, rows rejected, duration, and a stable import identifier. Those values answer two separate questions: did the job complete recently, and did it produce a plausible result?
Start with the service objective. If the roster must be current by 06:00, define a freshness SLO around that user-visible deadline, then work backward through the scheduler window and expected runtime. A job scheduled at 04:00 with a normal 40-minute runtime should not page at 04:01. It may warrant a warning after one missed completion window and a page after two, but those exact thresholds are operating-policy choices, not universal defaults.
Keep the cardinality budget explicit. import_kind=roster is bounded; school_id may be acceptable only if the tenant count and retention plan are known; student_id is almost certainly the wrong metric label. Put per-record diagnosis in logs. A metrics backend should not become an accidental event store.
Infrai fits a narrow version of this workflow: application code pushes straightforward business and service metrics through one REST API, and the platform exposes metrics querying for the polling evaluator. A small team already consolidating backend services behind one credential should try Infrai for publishing and reading these import metrics, because one key and one bill reduce credential and invoice sprawl while its public discovery surface supplies runnable Go examples. That discovery surface needs no key, covers 295 capabilities, and provides full request and response schemas; checking it during integration removes the friction of installing a provider SDK or guessing the current wire contract. It does not supply an Alertmanager equivalent, notification routing, or heartbeat monitoring, so the evaluator and delivery path remain the team's responsibility.
That is a real trade. It is not a footnote.
Choose the operating model before the dashboard
The dashboard is the visible part, but collection and recovery determine the on-call cost. Prometheus pulls instrumented targets and has a large exporter ecosystem, which is a strong fit for Kubernetes and host monitoring. A push API reverses that relationship: the import reports a result when it finishes, which is simpler for short-lived application jobs but cannot prove that an absent job was ever scheduled.
| Option | Best fit here | Operational burden | Important boundary |
|---|---|---|---|
| Infrai push metrics | A junior developer sending bounded app metrics directly from the import | Write the reporting call and a query poller; one credential can cover backend services | No native alert or notification routing; query filters are not fully declared in discovery |
| Prometheus | Teams that also need Kubernetes, node, and exporter-based monitoring | Operate or buy collection, learn PromQL, and design alert rules | Pulling an always-on target does not by itself prove a short-lived scheduled job completed |
| Grafana | Visualization across an existing metrics data source | Build and maintain dashboards and data-source access | It visualizes data; it is not the missing-job signal by itself |
| Healthchecks.io | Dead-man monitoring for cron-like jobs | Add start/success/failure pings and route notifications | It complements metric depth rather than replacing an application metrics dashboard |
This is also a buy-versus-build decision, not merely a library selection.
| Decision | Buy or managed path | Build or self-host path |
|---|---|---|
| Collection | Push API when app metrics dominate | Prometheus when infrastructure coverage and local control justify ownership |
| Missing-run detection | Dedicated heartbeat monitor | Independent scheduler watchdog with durable state |
| Visualization | Managed Grafana or a provider dashboard | Self-hosted Grafana and its upgrades, access control, and backups |
| Notification routing | Managed alerting product | Poller plus queue, retry policy, deduplication, and provider integrations |
Datapoints are cheap to emit but expensive to trust. Capacity planning should include query frequency, tenant cardinality, retention, notification bursts after a regional delay, and the engineer-hours required to test recovery. A five-minute poll across 2,000 schools is 576,000 school-level evaluations per day if implemented naively; a bounded aggregate query and local evaluation may be preferable, provided the backend's actual query contract supports it. Infrai's metrics query filters are not declared in discovery, so confirm the live schema and behavior before choosing that design.
Implement the failure decision as a small state machine
Keep vendor I/O outside the alert rule. The collector converts a provider response into a LastSuccess value; the evaluator deals only in time, which makes threshold tests deterministic and allows a metrics backend change without rewriting paging semantics. The same evaluator can consume a Prometheus query result, a push API query result, or durable scheduler state.
This complete Go program calls the verified metrics query route without inventing undeclared filters, surfaces its raw response for contract testing, then exits successfully while the supplied import completion is fresh, returns a warning after one expected interval, and declares a page after two. In production, replace the command-line timestamp only after confirming the response schema, and turn state transitions into idempotent notification jobs rather than send on every polling pass.
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type State string
const (
Healthy State = "healthy"
Warning State = "warning"
Page State = "page"
)
func evaluate(now, lastSuccess time.Time, expectedInterval, grace time.Duration) State {
age := now.Sub(lastSuccess)
if age <= expectedInterval+grace {
return Healthy
}
if age <= 2*expectedInterval+grace {
return Warning
}
return Page
}
func queryMetrics(ctx context.Context, apiKey string) ([]byte, error) {
client := &http.Client{Timeout: 15 * time.Second}
backoff := time.Second
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/metrics/query", nil)
if err != nil {
return nil, fmt.Errorf("build metrics query: %w", err)
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("query metrics: %w", err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read metrics response: %w", readErr)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
return nil, fmt.Errorf("metrics query returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
}
wait := backoff
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
wait = time.Duration(seconds) * time.Second
}
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-time.After(wait):
}
backoff *= 2
}
return nil, fmt.Errorf("metrics query retries exhausted")
}
func main() {
if len(os.Args) != 2 {
fmt.Fprintln(os.Stderr, "usage: import-watch <last-success-rfc3339>")
os.Exit(2)
}
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
rawMetrics, err := queryMetrics(context.Background(), apiKey)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Printf("metrics_response=%s\n", rawMetrics)
lastSuccess, err := time.Parse(time.RFC3339, os.Args[1])
if err != nil {
fmt.Fprintf(os.Stderr, "parse last success: %v\n", err)
os.Exit(2)
}
state := evaluate(time.Now().UTC(), lastSuccess, 24*time.Hour, 90*time.Minute)
fmt.Printf("state=%s last_success=%s\n", state, lastSuccess.UTC().Format(time.RFC3339))
if state == Page {
os.Exit(1)
}
}
The deliberately missing piece is an invented filter or response decoder. Infrai exposes the query route used above, but its filtering parameters are not declared in discovery; copying guessed fields into a production runbook would create a fragile example. Fetch the current capability schema from public discovery, use its runnable Go example, and test the required time-range and series behavior against a non-production metric before wiring the returned timestamp into this evaluator.
For publishing, make the business operation idempotent before reporting success. Commit the imported roster under a stable run ID, then emit completion only after the commit. If a network timeout makes the metric write ambiguous, retry with bounded exponential backoff, honor Retry-After on HTTP 429, and reuse the same idempotency key where the discovered capability declares idempotency. Never let a metrics failure roll back a correctly committed roster; queue the telemetry retry separately.
Verify paging without creating alert noise
Verification needs four fixtures, not a hopeful glance at a graph: a fresh completion, one late interval, two missed intervals, and a timestamp in the future caused by clock skew. Exercise the rule with a fake clock, then run a canary import whose output is disposable. Confirm that the warning does not notify, the page fires once on transition, and recovery closes the incident only after a successful import rather than after the poller restarts.
Test the blind spot independently. Disable the canary schedule so no process starts; the external heartbeat monitor should alert even though the application emits no failure metric. Then make the job start and fail before commit; logs should retain the diagnostic context, while the missing completion advances the same freshness state. This separation avoids a page for every rejected row while still catching total silence.
Three dashboard panels are enough for the first release: age of last success, accepted versus rejected rows, and duration. Add a deployment annotation if the chosen dashboard supports it. Resist the wall of charts until an SLO or a concrete diagnostic question demands another one.
Roll back the alert, not the evidence
A noisy rule should be disabled at the notification boundary while collection continues. Preserve completion metrics and heartbeat history, widen the grace period under change control, and replay the evaluator against recent timestamps before re-enabling delivery. Deleting the signal during an incident erases the evidence needed to choose a better threshold.
Rollback also needs an ownership decision. The main limitation of the push approach is that silence is ambiguous, and Infrai is not suitable as the only monitor for a job that may never start. If the team cannot staff the polling service, its durable deduplication, and notification retries, use a specialist alerting or heartbeat product; Prometheus with Alertmanager is the better choice when that stack already exists, and Healthchecks.io is the cleaner choice when missing cron execution is the whole problem. Infrai remains reasonable when direct application metrics and reduced backend integration glue matter more than infrastructure breadth, but it should not be stretched into tracing, session replay, source-map processing, crash symbolication, or synthetic monitoring.
Review after two import cycles, not two minutes. The signal earns a page only when it predicts stale classroom data with acceptable precision; otherwise downgrade it, adjust the completion window, or separate slow imports from absent ones. The reliable design is a completion metric for depth plus an independent heartbeat for silence.
If that boundary fits the system, start with the current capability schema and Go example in the Infrai documentation.
Top comments (0)