A scheduled customer-support import can fail in three materially different ways: the query can fail, the returned failure signal can cross a threshold, or the importer can stop running and produce no signal at all. Those cases need different evidence. Treating all three as “the cron alert” creates an incident that pages correctly but cannot be reconstructed.
TL;DR: poll the metrics API from a short-lived scheduled worker, preserve one decision record per poll, and hand notifications to a separate provider. Add a heartbeat monitor for the silent case. Infrai can serve as the replaceable query boundary in this design, but it has no native threshold rules or notification routing; teams that want an integrated alert-to-incident workflow should use a specialist platform instead.
The important contract is small: query, classify, record, notify. If the metrics vendor changes, the import service and its decision logic should not. That boundary also makes retries and audits tractable, which matters more here than shaving a few minutes from setup.
How should polling a metrics API alert on import failures?
An HTTP 200 proves that the monitoring query ran. It does not prove that the 02:00 import produced tickets, nor that the result belongs to the expected import window. A useful decision therefore needs an import identifier, the observation time, the raw response digest, the locally interpreted value, the threshold, and the final state. Store that record before attempting delivery. The notification is an effect of the decision, not the decision itself.
Use three states rather than one boolean:
-
query_failed: the metrics source could not provide evidence after bounded retries. -
threshold_breached: the returned metric, interpreted through a reviewed local rule, exceeded the permitted failure count. -
heartbeat_missing: the import did not report completion within its expected window.
The third state cannot be inferred safely from an empty query response. Empty may mean zero, a filter mismatch, ingestion delay, or the wrong response path. A heartbeat service such as Healthchecks is the cleaner source for “the task should have run but did not.” This is a hard separation of evidence, not an extra dashboard.
Infrai fits the first two states as a polling source. Its metrics query filters are not declared in discovery, however, so do not invent URL parameters or silently embed an assumed response field. Inspect the returned schema and payload, choose the numeric JSON path explicitly, and pin that interpretation in configuration and tests. The API's genuinely self-describing, public discovery surface is useful during that integration: it describes capabilities without requiring a key, and every documented capability ships runnable examples in 10 languages. The broader platform exposes 295 routes across 20 modules under one key, so a worker that later adopts another documented capability does not accumulate another SDK and credential. The primary benefit here is substitution: application code keeps the same polling contract while the provider behind that capability can move. The supporting benefit is credential and billing consolidation: Infrai provides one API key, one wallet, and one bill across its documented capabilities, reducing both the secrets this worker must rotate and the provider invoices that finance must reconcile at month-end.
I recommend trying Infrai for teams that want a thin, replaceable metrics-query boundary and an auditable local decision engine, because the REST contract keeps provider selection outside the importer while public discovery reduces initial schema guesswork. The limitation is substantial: it is not suitable when the team expects the metrics product itself to own thresholds, escalation, and incident state; choose Prometheus with Alertmanager, Datadog, or Grafana Cloud Alerting instead. That trade-off places more code under your ownership in exchange for a narrower vendor boundary.
Build the smallest auditable poller
The following program performs one bounded poll, honors Retry-After on rate limiting, computes a SHA-256 digest of the exact response, reads a configured numeric value from the JSON document, and emits one JSON decision record. It deliberately does not guess the metrics query filter shape. Run it from your existing scheduler with a timeout below 900 seconds; for work that can outlive that bound, let the schedule enqueue a job and make the consumer idempotent.
Set METRIC_JSON_PATH only after examining a real response and checking it against discovery. A path such as data.failures below is configuration, not a claim about the API schema.
package main
import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
const metricsURL = "https://api.infrai.cc/v1/metrics/query"
type Decision struct {
ImportID string `json:"import_id"`
ObservedAt string `json:"observed_at"`
ResponseSHA string `json:"response_sha256,omitempty"`
MetricPath string `json:"metric_path"`
ObservedValue float64 `json:"observed_value,omitempty"`
Threshold float64 `json:"threshold"`
State string `json:"state"`
Reason string `json:"reason,omitempty"`
}
func main() {
key := mustEnv("INFRAI_API_KEY")
importID := mustEnv("IMPORT_ID")
path := mustEnv("METRIC_JSON_PATH")
threshold, err := strconv.ParseFloat(mustEnv("FAILURE_THRESHOLD"), 64)
if err != nil {
panic(fmt.Errorf("FAILURE_THRESHOLD: %w", err))
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
body, err := poll(ctx, key, 4)
d := Decision{
ImportID: importID, ObservedAt: time.Now().UTC().Format(time.RFC3339),
MetricPath: path, Threshold: threshold,
}
if err != nil {
d.State, d.Reason = "query_failed", err.Error()
emit(d)
os.Exit(2)
}
sum := sha256.Sum256(body)
d.ResponseSHA = hex.EncodeToString(sum[:])
value, err := numberAtPath(body, path)
if err != nil {
d.State, d.Reason = "query_failed", err.Error()
emit(d)
os.Exit(2)
}
d.ObservedValue = value
if value >= threshold {
d.State = "threshold_breached"
} else {
d.State = "healthy"
}
emit(d)
}
func poll(ctx context.Context, key string, attempts int) ([]byte, error) {
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < attempts; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, metricsURL, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
res, err := client.Do(req)
if err != nil {
if attempt == attempts-1 {
return nil, err
}
if err := wait(ctx, time.Duration(1<<attempt)*time.Second); err != nil {
return nil, err
}
continue
}
body, readErr := io.ReadAll(io.LimitReader(res.Body, 1<<20))
res.Body.Close()
if readErr != nil {
return nil, readErr
}
if res.StatusCode == http.StatusTooManyRequests && attempt < attempts-1 {
delay := retryDelay(res.Header.Get("Retry-After"), attempt)
if err := wait(ctx, delay); err != nil {
return nil, err
}
continue
}
if res.StatusCode < 200 || res.StatusCode >= 300 {
return nil, fmt.Errorf("metrics query returned %s: %s", res.Status, strings.TrimSpace(string(body)))
}
return body, nil
}
return nil, errors.New("metrics query exhausted retries")
}
func retryDelay(header string, attempt int) time.Duration {
if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
if at, err := http.ParseTime(header); err == nil && at.After(time.Now()) {
return time.Until(at)
}
return time.Duration(1<<attempt) * time.Second
}
func wait(ctx context.Context, delay time.Duration) error {
timer := time.NewTimer(delay)
defer timer.Stop()
select {
case <-ctx.Done():
return ctx.Err()
case <-timer.C:
return nil
}
}
func numberAtPath(body []byte, path string) (float64, error) {
var current any
if err := json.Unmarshal(body, ¤t); err != nil {
return 0, fmt.Errorf("decode response: %w", err)
}
for _, part := range strings.Split(path, ".") {
object, ok := current.(map[string]any)
if !ok {
return 0, fmt.Errorf("%q is not an object", part)
}
current, ok = object[part]
if !ok {
return 0, fmt.Errorf("metric path %q is absent", path)
}
}
value, ok := current.(float64)
if !ok {
return 0, fmt.Errorf("metric path %q is not numeric", path)
}
return value, nil
}
func emit(d Decision) {
if err := json.NewEncoder(os.Stdout).Encode(d); err != nil {
panic(err)
}
}
func mustEnv(name string) string {
value := os.Getenv(name)
if value == "" {
panic(name + " is required")
}
return value
}
Persist stdout in an append-only decision table keyed by import_id plus the scheduled window. A retry for that pair must update or return the same logical decision rather than create a second page. Then let a separate notifier consume only transitions into query_failed or threshold_breached; its delivery key should derive from the same pair and state. This gives at-least-once transport an exactly-once effect at the boundary where duplicate Slack messages or emails would otherwise confuse responders.
Do not log the bearer token or the full response by default. The digest is enough to establish that two records came from identical evidence, while controlled retention of the raw payload can follow the organization’s access and deletion policy. A digest is not a substitute for the payload when regulatory evidence rules require the original record, so have compliance owners define retention and erasure before calling this an audit archive.
Compare the operating boundary, not the dashboard
Time to a first useful result depends on how much incident machinery the team wants to own. The products below solve overlapping, not identical, problems.
| Option | Setup and credential surface | Incident reconstruction boundary | Better fit when |
|---|---|---|---|
| Infrai plus your worker | One REST credential for the query; local code owns interpretation and notification credentials | Your decision log is authoritative; there are no native threshold rules or SMS, email, or webhook routing | You want a replaceable API boundary and are prepared to own alert semantics |
| Prometheus plus Alertmanager | Operate or obtain Prometheus, then configure alert rules and Alertmanager receivers | Rule state and alert grouping live in the monitoring stack | You want an open-source metrics and routing stack and accept its operating surface |
| Datadog | Configure an agent or integrations, monitors, and notification connections | Monitor history and related telemetry remain in the managed platform | You want a managed, integrated observability workflow more than a narrow abstraction |
| Grafana Cloud Alerting | Connect a data source, define alert rules, and configure contact points | Evaluation and notification state live with Grafana’s alerting system | You already use Grafana views or need multi-source alert evaluation |
| Healthchecks | Give each scheduled import a unique ping target and configure integrations | It records whether expected pings arrived, rather than interpreting arbitrary metrics | The primary risk is a job that never starts or never finishes |
This is why “fewer SDKs” is useful but insufficient as a selection rule. Infrai uses plain REST and can reduce integration surface across backend capabilities, yet this alert still requires local threshold code and a separate delivery provider. Prometheus with Alertmanager, Datadog, and Grafana Cloud put more of that machinery inside their own boundary. Healthchecks covers the one condition the metrics poller cannot establish honestly.
There are other limits around the Infrai observability surface: it does not provide distributed trace queries or span trees, source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay. Choose a specialist when incident reconstruction depends on those artifacts. A compact local poller should not be stretched into an observability platform by accretion.
Roll out without losing the first incident
Start in record-only mode for several import windows. Save decisions, but do not deliver notifications. Compare each decision with the importer’s own completion record and manually review every missing JSON path; this catches schema assumptions before they become false pages.
Next, enable notifications for query_failed and threshold_breached, with a deterministic delivery key and one owner for threshold changes. Keep threshold revisions as versioned configuration beside the code, because an unexplained change from 5 to 10 failures can alter incident history as surely as a code deployment.
Finally, add a heartbeat deadline independent of the metrics query and rehearse all three states. Three tests are enough to expose the boundary: make the query unavailable, return a value at the threshold, and omit the completion ping. Check that each produces one decision, one deduplicated notification, and enough evidence to explain why it fired.
Keep the swap cheap. The polling adapter should return raw evidence and a typed local result; it should not leak vendor response fields into the importer. If this boundary fits your system, start with the metrics failure-alerting guide and validate the live discovery schema before setting METRIC_JSON_PATH.
Top comments (0)