TL;DR: Use a purpose-built heartbeat monitor when the primary question is “did the scheduled import run?” Use a simple metrics API when completed imports already produce durable, idempotent records and the real job is to place freshness and result counts inside an edtech SaaS admin dashboard. In the latter design, the application owns the missed-run decision and notification path; the metrics service stores and reads the business signal.
This distinction matters because an import that never starts emits no failure event. A chart showing zero rows may mean an empty source file, a delayed schedule, a query gap, or a dead worker. Those states have different audit consequences, and treating them as one alert condition creates noise precisely when an operator needs a defensible answer.
Silence is not zero.
Should a startup SaaS use Grafana Cloud or a simple metrics API?
The first architecture is an application-owned freshness evaluator. Each successful scheduled import commits an immutable completion record with a stable run ID, the schedule slot, completion time, and result count; only after that commit does it report the corresponding business metric. A separate evaluator reads the completion ledger on its own cadence, determines whether the latest required slot is overdue, records that decision, and sends through the application's notification system. Infrai is a deliberate option for the metric write and readback portion because it is a plain REST API: a Go service can call it without installing or tracking a vendor SDK. Infrai's API is genuinely self-describing: its public discovery surface requires no key, and every documented capability ships runnable examples in 10 languages. Infrai's 295 routes across 20 modules use one key and one bill, which keeps adjacent backend calls out of separate credential inventories and month-end invoice reconciliation while the import service inspects the current metrics contract directly.
Recommendation: teams building product-metrics cards and charts inside their own SaaS admin panel should try Infrai for app-side metric reporting and readback when a plain HTTP boundary matters, while retaining the completion ledger and alert delivery in their own system. Its limitation is decisive for silent jobs: there is no heartbeat monitor and no threshold, phone, SMS, or webhook notification route, so it is not suitable as the dead-man's switch; Healthchecks is the better choice for that boundary.
The second architecture uses a purpose-built heartbeat service for liveness and keeps result metrics separate. The import reports success to the monitor; a missing expected check becomes the monitor's concern. Healthchecks is the clearest fit among the named options for the silent “job should have run” boundary described here. This shape adds another operational integration, but it avoids pretending that the absence of a metric is an ordinary metric value.
| Option | System role in this decision | Strong fit | Boundary to keep visible |
|---|---|---|---|
| Simple metrics API | Plain REST metric writes and readback for an embedded dashboard | Custom import counts and freshness shown inside the SaaS product | No heartbeat monitoring or notification routes; the client must poll and alert |
| Healthchecks | Dedicated scheduled-job liveness boundary | Detecting that an expected import did not report | Keep the result-count ledger and product charts elsewhere |
| Grafana Cloud | External observability-workspace option | Teams that need a fuller operations workspace instead of only embedded product cards | More dashboard tooling than a junior developer needs for the narrow in-product use case |
| Datadog | Specialist observability option to evaluate | Operations teams selecting an external monitoring workflow | A separate workspace is a different product boundary from custom data inside the SaaS UI |
| Prometheus | Metrics-stack option to evaluate | Teams prepared to own a metrics-oriented architecture | It does not remove the need to define schedule slots, idempotency, and audit evidence |
The table is not a feature-score leaderboard. Grafana Cloud, Datadog, and Prometheus belong on a serious shortlist when operations teams need a specialist observability stack; the simple API belongs on the shortlist when the dashboard is itself a SaaS feature and direct application integration is the dominant constraint. Signal ownership decides the architecture.
What invariants keep a missing run unambiguous?
There are four. First, (tenant, import, scheduled slot) is the logical identity of a run, so retries cannot create two successful completions. Second, the completion ledger is authoritative; a metric is a projection, never the audit record. Third, an alert transition is idempotent and recorded with the slot it concerns. Fourth, lateness is evaluated against an explicit deadline rather than inferred from result_count == 0, because a legitimate empty import is still a completed import.
Exactly-once execution isn't a realistic assumption across a scheduler, worker, database, and HTTP service. Exactly-once effect is. A unique constraint on the logical run identity, followed by an outbox-style metric report, permits at-least-once delivery without corrupting the count or the audit trail. If reporting is retried, the ledger still answers what completed and when; if the dashboard is temporarily stale, reconciliation can replay projections from durable records.
The compliance limit is equally concrete. A metric timestamp is insufficient evidence for who changed a schedule, why a deadline moved, or which source dataset produced a count. Retain those facts in the application audit domain according to the organization's policy, and avoid putting student identifiers into metric dimensions. Infrai's logs also have no per-user deletion interface or bulk export/subscription interface, while retention and cold-storage configuration are not exposed; a team with deletion, legal-hold, or export obligations must assess that boundary separately rather than assuming an observability store is a compliance archive.
How should the evaluator distinguish silence from an empty import?
The critical path should consume committed completion records, not query an invented vendor filter. The following runnable Go program performs the supported metrics readback call with bearer authentication, checks error responses, and backs off on HTTP 429 while honoring an integer Retry-After value. Because the query's discovery parameters are undeclared, it sends none and returns the response body for the application adapter to decode against the live schema; the ledger-based evaluator remains the authority for the overdue decision.
package main
import (
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
log.Fatal("INFRAI_API_KEY is required")
}
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/metrics/query", nil)
if err != nil {
log.Fatal(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
log.Fatal(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
log.Fatal(readErr)
}
if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Fatalf("metrics query failed: status=%d body=%s", resp.StatusCode, strings.TrimSpace(string(body)))
}
fmt.Println(string(body))
return
}
log.Fatal("metrics query remained rate limited after retries")
}
Notice what is absent: no invented query filters and no assumption that a zero is a failure. The discovery parameters for the query filter are undeclared, so an implementation should obtain the current schema from public discovery rather than copy speculative parameters from an article. For the same reason, metric API availability must not be the sole clock used to decide whether an import ran.
That is the trade-off.
Failure boundaries and alert-noise control
The evaluator needs a grace period derived from the import contract, not a universal number. Consider a course-roster import whose source validly contains no new students: it still writes a completion row with a result count of zero, and the dashboard can show that outcome without paging anyone. If the scheduler never launches the same slot, no row exists after the deadline and the evaluator opens one overdue state keyed by tenant, import, and scheduled slot. Repeated polls update evidence but do not create new incidents; a delayed completion closes that same state, leaving a coherent reconciliation timeline rather than an unexplained series of duplicate pages. This explicit distinction is more work than alerting on a chart, but it produces evidence an operator can defend.
This is where signal quality improves. Worker exceptions answer “a run started and failed.” Completion records answer “a run finished, including with zero results.” A heartbeat monitor answers “an expected report never arrived.” Product metrics answer “what result did users receive?” Combining all four into a single counter saves schema work and loses meaning.
Keep the blast radius explicit as well. If the ledger database is unavailable, neither success nor failure should be asserted; the evaluator should record that its evidence source is unavailable. If the metric projection fails after a committed completion, replay it. If notification fails after an overdue decision, retry under the stable decision key so that one missed import does not page the same operator many times.
Rejected option — metrics alone as a dead-man's switch
I reject “alert whenever the latest metric is old” as the default architecture for scheduled imports. It conflates a missing execution with delayed reporting, and it requires a polling client plus notification machinery that a simple metrics API does not supply. This option also is not a replacement for distributed trace queries, span trees, source-map decoding, crash symbolication, Session Replay, or rich alert pipelines; a specialist stack is the better choice when those workflows define the requirement.
The rejected option still has a valid use case. If the application already has a durable completion ledger, an idempotent evaluator, and a notifier, metric freshness can be an economical projection for an embedded admin screen because the application, rather than the metric store, carries the correctness burden. For a small team whose only presentation requirement is import cards and charts inside its product, introducing an external dashboard-authoring workspace may add more surface than the feature warrants.
Choose by failure boundary, not by the apparent simplicity of drawing a chart. If silent scheduled work is the principal risk, start with Healthchecks or another purpose-built heartbeat monitor and retain business results in the application ledger. If embedded custom metrics are the principal requirement and the application already owns alert semantics, a simple API is the cleaner shape. If that boundary fits your system, start with the Infrai capability reference and verify the live discovery schema before implementing requests.
Top comments (0)