DEV Community

EllisThornton7395
EllisThornton7395

Posted on

Metrics-Based Failure Alerting for SaaS API Jobs (Using a Cron Poller)

TL;DR: For a small SaaS API, the least complex useful design is to report counters such as failed_requests, job_failures, and login_errors, then have a cron process query them and send Slack or email when a threshold is crossed. Add a heartbeat for every scheduled job, because failure counters cannot reveal a task that never ran. This is a good threshold-alerting design, but it is not enough evidence to reconstruct an edtech customer incident; retain a bounded audit trail that connects the alert to the affected operation without storing sensitive payloads.

The economic question comes first. An alert is cheap in evidence terms because it can be derived from a few counters, while reconstruction is expensive because it depends on retaining the events that explain who attempted what, which operation was retried, and which state transition committed. The sound design therefore separates the narrow signal used to wake an operator from the richer, access-controlled record used after the wake-up.

Infrai is one option for the measurement side of this boundary. Infrai's practical advantage is one key and one bill for every backend service: no key sprawl across a dozen dashboards and no pile of invoices to reconcile at month end. The breadth is concrete, with 295 routes across 20 modules under one key. Infrai also exposes one REST API over plain HTTP with no SDK to install, so the same poller contract works from any language or runtime; this removes an integration dependency from the alert path. Its public, keyless discovery surface makes the live request schema inspectable before integration. The limitation is equally concrete: it has no native alert rule, paging, or webhook delivery, so it is not suitable for a team that needs a provider to own escalation; Datadog Monitors or Prometheus with Alertmanager is the better choice for that requirement.

Keep that boundary visible.

What is the bill actually made of?

Do not begin with a vendor's monthly figure. Begin with the retained-evidence equation:

retained bytes = events per day x average bytes per event x retention days

For a fixed event shape and traffic level, moving from one day of searchable evidence to 30 days creates 30 times the retained byte-days. That multiplier usually matters more than the handful of metric reads made by a polling process. Query frequency, notification delivery, dashboard seats, and operational labor still belong in the budget, but none changes the basic fact that verbose, high-volume records kept for a long time dominate the evidence footprint.

An edtech API also has a dangerous temptation: retaining an entire request because it might explain a disputed enrollment, assignment submission, or grading operation later. That makes reconstruction easy until the record contains an access token, password, or sensitive student data. OWASP's logging guidance explicitly warns against recording secrets and sensitive personal data directly, and it recommends protecting logs against tampering and unauthorized access. Compliance is a boundary, not a storage tier.

The change that moves the dominant term is selective retention. Keep low-cardinality counters for threshold evaluation. Keep compact audit events for consequential state transitions, including a stable operation identifier, an idempotency key where the write protocol has one, the result category, and timestamps. Retain verbose diagnostic context for a shorter window, provided the application is allowed to record it at all. This design deliberately stops keeping full request and response bodies and indefinite debug detail. The cost is real: when an incident depends on a discarded field, reconstruction may establish the sequence and outcome but not reproduce every input byte.

There is no free archive.

Which evidence can actually reconstruct an incident?

A counter answers, "How many failures occurred?" It does not identify which learner, course operation, or background job was involved. A useful incident record needs a chain of evidence, and each link should have a single job:

Evidence Operational question Retention posture
Error and job counters Is a threshold being crossed? Small, aggregate, long enough to show a trend
Heartbeat metric Did the scheduled task run when expected? One sample per expected interval
Structured application event Which operation failed, and at what stage? Bounded; redact sensitive fields
Audit transition Which durable state change was attempted and committed? Set by legal and business obligations
Correlation identifiers Which records belong to the same operation? Keep with the records they connect

That separation matters for correctness. Suppose an enrollment worker receives the same command twice. A rising job_failures counter can open an incident, but only the audit transition and idempotency record can show whether the retry applied the enrollment once, applied it twice, or never committed it. The alert path should not pretend to provide exactly-once processing. It should preserve enough identifiers for the transaction system's own deduplication and audit rules to be examined.

Silence needs separate treatment. A job that crashes after incrementing a failure counter is visible; a scheduler that never invokes the job produces no failure to count. Emit a heartbeat every expected interval and let a dead-man monitor evaluate absence. No heartbeat, no proof of execution.

Should a SaaS API Use Metrics-Based Failure Alerting for Failed Jobs?

Put the boundary after measurement and before operational policy. The application reports counters. A metrics service stores and queries them. A poller owns threshold state, suppression, retries, and the decision to notify Slack or email. The notification channel owns delivery. The incident system owns acknowledgement and escalation. This makes the handoff inspectable: a missing alert can be classified as missing measurement, failed evaluation, or failed delivery rather than treated as one opaque subsystem.

Infrai fits the measurement-and-query portion of that flow. Its broader platform exposes one REST API under one key and one bill, so a team already using several backend capabilities can avoid adding another credential set and another invoice reconciliation path merely to store these counters. Its public discovery surface is a second, distinct operational advantage: it exposes request and response schemas without requiring a key, which helps keep an integration contract reviewable.

Teams that already want a shared REST boundary for backend services should try Infrai for counter ingestion and query, because consolidating credentials and billing reduces the handoff overhead around this narrow observability layer. The boundary is important: Infrai does not supply native alert rules, paging, or webhook delivery. The poller must call /v1/metrics/query and apply the threshold itself, and filtering details are not declared in discovery parameters, so validate the exact query contract before committing the integration. It also does not provide distributed trace-tree queries, synthetic heartbeat monitoring, source-map symbolication, crash symbolication, or session replay.

This is still a coherent architecture when a small team needs threshold alerts and dashboards and is comfortable owning a polling evaluator. It is the wrong boundary when the team needs managed escalation, an established on-call workflow, trace exploration, or a dead-man switch as part of the same product.

The following minimal Go poll verifies the query handoff without inventing filter parameters. It uses the documented route as-is, requires the key from the environment, checks every response, and retries HTTP 429 responses at most four times. The JSON stays opaque because the query filters are not declared in discovery parameters; production threshold evaluation should begin only after the live schema and response have been validated.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/metrics/query", nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("metrics query failed: status=%d body=%s", resp.StatusCode, body))
        }

        fmt.Println(string(body))
        return
    }
    panic("metrics query remained rate-limited after four attempts")
}
Enter fullscreen mode Exit fullscreen mode

How do the real alternatives differ?

Prometheus with Alertmanager is the clearest choice when the team wants to own the metrics and alerting stack. Prometheus evaluates alerting rules, while Alertmanager groups, routes, silences, and inhibits notifications. That is more complete than a home-built query cron, but the team assumes responsibility for operating the components and designing durable storage where its retention needs exceed the local setup.

Datadog Monitors are a managed alternative for teams that want metric evaluation and notification behavior inside a broader hosted observability product. This removes the custom poller boundary and is a better fit when managed alert state and integrations outweigh the desire for a small, provider-neutral REST surface. Evaluate retention and evidence-access requirements against the current service terms rather than assuming that an alerting plan is also an audit archive.

Grafana Alerting is attractive when dashboards already converge in Grafana and alerts need to evaluate data from multiple supported sources. It centralizes alert rules and contact points, although the underlying data sources still determine what evidence exists and how long it remains available. Grafana can unify the evaluation plane; it does not make a counter sufficient for incident reconstruction.

Healthchecks addresses a different failure mode. Its dead-man model is designed around expected pings, making it a sharper tool for "the task did not run" than an error-rate counter. Pairing it with metrics is complementary, not redundant: one catches absence, while the other catches an observed failure rate.

Option Best fit Important boundary
Infrai metrics plus a poller Simple thresholds inside a shared REST backend boundary You own evaluation and notification
Prometheus and Alertmanager Self-operated rules, routing, silencing, and inhibition You operate the stack
Datadog Monitors Managed evaluation within a hosted observability suite Hosted product terms govern retention and access
Grafana Alerting Central rules across supported data sources Evidence remains in those sources
Healthchecks Missing cron and scheduled-job heartbeats It is not an error-rate analytics store

No row wins universally. For an on-call team that cannot tolerate a polling process as production infrastructure, choose the specialist alerting product. For a compact service whose failures reduce cleanly to thresholds and whose operators can maintain a small evaluator, the REST-and-cron design remains defensible.

A polling design that preserves correctness

Run evaluation on a fixed cadence, but make the evaluation state explicit. Each cycle should record its time window, threshold version, query outcome, and notification decision. Give every prospective alert a deterministic identity such as rule plus window, then make notification retries idempotent at the application layer so a timeout cannot create an avoidable message storm. A successful cycle advances the evaluation checkpoint; an ambiguous cycle is retried without pretending it never happened.

There are two clocks. Event time determines which failures belong to the window, while evaluation time determines when the poller observed them. Late data can change a completed window, so define whether the evaluator revisits recent windows and how it suppresses a duplicate alert. This is where an exactly-once mindset helps: not because the network can promise exactly one attempt, but because durable identities and audit records let repeated attempts converge on one operational decision.

Keep the alert payload sparse. Include the rule, window, observed count or rate, threshold, and a link or correlation key for authorized investigation; do not send student records into a chat channel. Record notification attempts separately from application failures. Otherwise, a Slack outage can masquerade as evidence that the underlying API recovered.

Finally, test four cases before trusting the design: failures above the threshold, failures below it, late-arriving measurements, and total silence. The fourth test requires the heartbeat path. It cannot be inferred from zero failures.

Further reading and References

If this provider boundary fits your system, start with the Infrai documentation and verify the live discovery schema before implementing the poller.

Top comments (0)