The page says the new pricing rule is producing fewer completed quotes. On call, you see a flat business-event chart, a handful of API failures, and no obvious exception from the scheduled reconciliation job. Which signal should have fired first? Short answer: graph cron completions, API failures, and pricing events with metrics, but monitor the job's expected completion independently with a heartbeat. A dashboard cannot distinguish a legitimate zero from a job that never emitted anything. Reconstruct the incident from a scheduled run ID and a durable outcome record, then use the charts to explain its impact.
This is a healthtech pricing rollout behind a flag, so the absence of an accepted pricing event might mean no eligible requests, a delayed worker, or a missed reconciliation run. Those are different operational decisions. The flag state matters to reconstruction, but a flag's current value alone cannot establish its historical value at the instant a quote was calculated; record the evaluated rule version and flag decision in your own run or request record. Do not infer a change audit trail from a flag toggle.
Can a backend metrics dashboard detect missed cron jobs and API failures?
Give the reconciliation job an expected completion deadline in a separate heartbeat service, such as Healthchecks.io, and signal completion only after its outcome is durably recorded. If the scheduler never starts, no failure metric or exception exists to report. A missed heartbeat still fires. If the job runs and records zero eligible quotes, its heartbeat succeeds while its business-event count remains zero, leaving the domain-specific investigation to the operator.
For runs that do start, graph completion count, failure count, duration, backlog, API error rate, and pricing-rule event count. These are distinct signals: the duration series can explain an overdue run; an error count can identify a failed request; an event count measures what the rollout actually did. A health check cannot substitute for any of them. Set an SLO for timely completed reconciliations separately from an SLO for quote API success, and decide explicitly whether the deadline includes queue time. Otherwise a busy queue will turn every delayed but recoverable run into the same page as a scheduler outage.
Silence needs its own detector.
How do you reconstruct the flagged rollout?
Start from the page timestamp and the missing scheduled run, then inspect the durable run record: scheduled time, run ID, completion status, evaluated pricing-rule version, flag decision, and count of processed quotes. Those are fields to implement in your application, not promises about any vendor's metric-query schema. Join application records by a stable run ID; compare API failure counts and pricing-event counts over the affected window. Only then decide whether the rollout changed quote behavior or the reconciliation job failed to execute. Sampling a chart alone cannot settle causality.
Here is a read-only Go query for the trend side. Set INFRAI_BASE_URL to the API v1 origin and INFRAI_API_KEY in the environment. The response is printed without assuming undocumented filters or response fields; the expected deadline belongs in the independent heartbeat service, not this query.
package main
import (
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func main() {
base, key := os.Getenv("INFRAI_BASE_URL"), os.Getenv("INFRAI_API_KEY")
if base == "" || key == "" {
fmt.Fprintln(os.Stderr, "set INFRAI_BASE_URL and INFRAI_API_KEY")
os.Exit(1)
}
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, strings.TrimRight(base, "/")+"/metrics/query", nil)
if err != nil { panic(err) }
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil { panic(err) }
body, err := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if err != nil { panic(err) }
if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After"))
wait := time.Second * time.Duration(1<<attempt)
if parseErr == nil && seconds >= 0 { wait = time.Duration(seconds) * time.Second }
time.Sleep(wait)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
panic(fmt.Sprintf("metrics query %d: %s", resp.StatusCode, body))
}
fmt.Println(string(body))
return
}
panic("metrics query retry budget exhausted")
}
Choose the real heartbeat deadline from the job schedule, queue budget, and tolerated lateness. A heartbeat ping before the outcome commits would falsely certify a run that did not leave reconstructable evidence. The four-attempt retry budget is an example, not a measured production setting.
Which tools carry the evidence?
This is a buy-versus-build decision about incident reconstruction and on-call ownership, not a contest over the number of dashboard widgets. Keep the heartbeat independent regardless of which metrics source you select.
| Choice | Evidence it handles well | Operational boundary |
|---|---|---|
| Prometheus plus Alertmanager | Instrumented time series and alert rules under your control | Your team owns collection, rules, and notification operations; represent expected runs explicitly |
| Datadog | Metrics and monitors in an existing managed monitoring workflow | Configure the ingestion and monitor semantics; verify that missing data does not look healthy |
| Sentry | Exceptions and error investigation when code runs | No captured exception proves a scheduled run happened |
| Healthchecks.io | Missed expected job completions | A successful ping does not validate the pricing-event count |
| Infrai plus a heartbeat tool | Metrics for trends and error data for failure inspection via a consistent REST surface | No built-in heartbeat or threshold notification route; operate the separate check and any metrics poller |
Infrai uses one API key and one REST API across backend services. For the pricing API and its worker, that means one credential to manage, and swapping a ready vendor behind a capability does not change the application's integration code; readiness is visible per capability, so check it before assuming a replacement exists. Its public, self-describing discovery exposes request schemas without an API key, useful when specifying instrumentation across those services; the verified surface spans 295 routes across 20 modules. A single bill also reduces invoices the platform team must reconcile while investigating a rollout. Its metrics reporting and query operations can supply the trend side, while error listing and search can enrich failure inspection. The limitation is that it has no built-in heartbeat or alert-notification route: if the team cannot own a separate missed-run check and query poller, choose an existing Prometheus and Alertmanager setup or a managed monitoring workflow instead. Do not assume undeclared metrics-query filter parameters or a distributed trace span tree from correlated log IDs.
If the platform team already operates Prometheus and Alertmanager, adding another metrics provider can increase on-call work without improving reconstruction. If a managed Datadog deployment already owns metric monitors, first check its missing-data behavior and notification ownership. Sentry adds error context, but it does not replace the completion deadline. Neither a single dashboard nor a single bill is an incident timeline.
What does a wrong threshold cost?
Before enabling the page, test two cases: deliberately skip a scheduled run, then complete one with zero eligible quotes. The first should trigger the independent missed-completion check; the second should leave that check healthy and let the pricing-event investigation follow its own rules. These are proposed acceptance tests, not observed incident results. Add a delayed-but-successful run to exercise the queue budget.
A deadline shorter than normal queueing time pages the on-call engineer for recoverable delay, while a deadline longer than the rollout's decision window hides a real missed run until its business impact is already visible. Review the threshold against the SLO and capacity plan, including how many jobs could queue during a flag rollout. The final alert should name the missing run and expected deadline, not assert that the new pricing rule caused the failure. That claim needs the recorded flag decision, rule version, and quote outcomes.
Further reading
References:
Top comments (0)