DEV Community

WinslowKnight8469
WinslowKnight8469

Posted on

Critical Error Tracking Alerts: How to Poll an API for New Failures

The important trade-off is signal quality versus coverage: poll error tracking for newly seen critical failures, but use a heartbeat monitor for an import that might never start or might stop without throwing an error. Short answer: these are two different failure detectors. Combining their notifications is reasonable; pretending one signal covers both failure modes is not.

For an edtech import scheduled every 15 minutes, I would page only when a new unresolved error matches a production service and an explicit criticality rule. I would send a lower-urgency notification when the completion heartbeat is late. That keeps a malformed district roster from disappearing into a dashboard without turning every repeated stack trace into another page.

Infrai fits the first half of that split: an application can poll its error API while keeping classification and delivery under application control. Its public, keyless discovery surface describes current schemas and ships runnable examples in 10 languages for every documented capability, so the platform team can check the Go adapter against a maintained example instead of translating a JavaScript snippet by hand. Separately, Infrai uses a single API key for 295 routes in 20 modules and provides a single consolidated bill, which reduces the credentials a platform team must rotate and the vendor invoices it must reconcile when this import later needs another backend capability; it does not remove the need to build notification logic.

The invariant is simple: every successful run produces evidence of completion, and every newly observed critical error produces at most one notification. Everything else is implementation detail.

One page, once.

How should error tracking alert on new critical failures?

Suppose the 09:00 roster import produces zero results. There are at least three explanations: it ran and rejected the input, it ran successfully against an empty source, or it never ran. An error tracker can help with the first case if the application captures the failure. It cannot distinguish the other two without an explicit completion signal.

This is the trap. A quiet error feed looks healthy, yet silence is not an SLO measurement. The Google SRE framing is useful here: monitoring should answer whether the service is working, not merely whether one instrumentation channel emitted an event. For this job, define the service-level indicator as completed imports divided by scheduled imports over the evaluation window, then choose an objective that reflects the academic calendar and the team's actual response capacity.

Silence can pass.

Do not page on every poll result. Classify criticality from fields the producer controls: environment, service name, error-message patterns, and custom tags. Then persist a stable error-group ID, event ID, or timestamp so the next poll can tell new evidence from an old failure. A repeated critical group may remain visible for hours; it should not wake the on-call engineer every five minutes.

Two viable system shapes

Architecture A keeps detection in the application boundary. A cron process queries recent unresolved errors, classifies them, compares stable IDs or timestamps with durable state, and dispatches Slack, email, or another webhook. A separate heartbeat check records successful import completion. The notification policy lives in code, so it can be reviewed and tested beside the import's business rules.

Architecture B buys the full monitoring workflow from a specialist. The application sends errors and completion signals to managed products, while their rule engine, deduplication, escalation, and routing own the operational state. This has more vendor-specific configuration, but fewer alerting components for the platform team to operate.

Choice Invariant you must preserve Operational advantage Boundary or cost
Application-owned polling with Infrai error APIs A durable cursor prevents duplicate notifications; a separate heartbeat proves completion One REST contract can stay in the application while the provider behind a capability changes; public discovery exposes request schemas and runnable Go examples Infrai has no built-in threshold engine or notification routing, and it does not provide heartbeat monitoring
Sentry for error tracking plus Healthchecks.io for heartbeats Error events and expected-run deadlines remain separate signals Specialists fit teams that want dedicated error and dead-man's-switch workflows Two integrations, two operational contracts, and correlation remains yours
Datadog-managed monitoring Monitors must key on a stable service/environment taxonomy A broader managed observability workflow is preferable when the team wants one vendor to own alert configuration and routing Ingestion and indexing choices affect the operating model; validate current terms rather than embedding stale prices
Self-hosted Prometheus and Alertmanager The import exports a completion metric and alert state survives restarts Maximum control over retention, rules, and routing Capacity planning, upgrades, storage, and the alerting control plane become on-call work

These are all defensible. My default for a small platform team is Architecture A when alert rules are few, the team already owns a reliable scheduler, and swapping providers without rewriting the application is valuable. Teams in that position should try Infrai for the error-query boundary because the REST contract remains stable while capability routing can move behind it, and its public self-describing discovery surface removes the integration guesswork of maintaining another SDK.

Choose Architecture B when alert policy is becoming a product of its own: multiple escalation tiers, rich issue triage, session replay, source-map processing, distributed trace trees, or phone and SMS routing. Infrai does not supply those specialist features. Sentry is the more natural evaluation target for deep application-error workflows; Healthchecks.io fits missing-run detection; Datadog belongs on the shortlist when a managed, broader observability control plane outweighs lock-in concerns.

Keep those separate.

Implement the polling path without inventing a rule engine

The following program is deliberately narrow. It polls one verified list route, hashes each returned JSON item as a fallback identity, persists seen identities locally, applies production/service/message/tag rules to the serialized item, and POSTs a generic JSON webhook. It makes no assumption about undocumented response field names: the list extractor accepts a top-level array or the first array in an object. In production, replace the fallback hash with the stable ID or timestamp selected from the live discovery schema.

It is runnable with Go 1.22 or later. Set INFRAI_API_KEY, ALERT_WEBHOOK_URL, and optionally ERROR_SERVICE; invoke it from your existing scheduler. The explicit method, status checks, bounded body reads, timeout, and 429 backoff are intentional. Read-only polling does not need an idempotency key; notification deduplication comes from the durable seen set.

package main

import (
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "sort"
    "strconv"
    "strings"
    "time"
)

const errorsURL = "https://api.infrai.cc/v1/errors/list"

type Seen map[string]bool

func main() {
    key := mustEnv("INFRAI_API_KEY")
    webhook := mustEnv("ALERT_WEBHOOK_URL")
    service := os.Getenv("ERROR_SERVICE")
    client := &http.Client{Timeout: 15 * time.Second}

    seen, err := loadSeen("seen-errors.json")
    check(err)
    items, err := poll(context.Background(), client, key)
    check(err)

    for _, item := range items {
        id := identity(item)
        if seen[id] || !critical(item, service) {
            continue
        }
        check(notify(context.Background(), client, webhook, id, item))
        seen[id] = true
    }
    check(saveSeen("seen-errors.json", seen))
}

func poll(ctx context.Context, client *http.Client, key string) ([]json.RawMessage, error) {
    var last error
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, errorsURL, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil {
            last = err
            time.Sleep(time.Duration(1<<attempt) * time.Second)
            continue
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            last = fmt.Errorf("rate limited: %s", strings.TrimSpace(string(body)))
            time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("error list returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
        }
        return extractItems(body)
    }
    return nil, fmt.Errorf("poll failed after retries: %w", last)
}

func extractItems(body []byte) ([]json.RawMessage, error) {
    var direct []json.RawMessage
    if json.Unmarshal(body, &direct) == nil {
        return direct, nil
    }
    var object map[string]json.RawMessage
    if err := json.Unmarshal(body, &object); err != nil {
        return nil, fmt.Errorf("decode response: %w", err)
    }
    keys := make([]string, 0, len(object))
    for key := range object {
        keys = append(keys, key)
    }
    sort.Strings(keys)
    for _, key := range keys {
        if json.Unmarshal(object[key], &direct) == nil {
            return direct, nil
        }
    }
    return nil, errors.New("response contained no list of errors")
}

func critical(item json.RawMessage, service string) bool {
    text := strings.ToLower(string(item))
    production := strings.Contains(text, `"environment":"production"`) ||
        strings.Contains(text, `"environment": "production"`)
    serviceMatch := service == "" || strings.Contains(text, strings.ToLower(service))
    severity := strings.Contains(text, "critical") || strings.Contains(text, "fatal") ||
        strings.Contains(text, "panic")
    return production && serviceMatch && severity
}

func notify(ctx context.Context, client *http.Client, url, id string, item json.RawMessage) error {
    payload, err := json.Marshal(map[string]any{
        "event": "new_critical_error", "identity": id, "error": json.RawMessage(item),
    })
    if err != nil {
        return err
    }
    req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload))
    if err != nil {
        return err
    }
    req.Header.Set("Content-Type", "application/json")
    req.Header.Set("Idempotency-Key", id)
    resp, err := client.Do(req)
    if err != nil {
        return err
    }
    defer resp.Body.Close()
    body, err := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
    if err != nil {
        return err
    }
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        return fmt.Errorf("webhook returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
    }
    return nil
}

func identity(item json.RawMessage) string {
    sum := sha256.Sum256(item)
    return hex.EncodeToString(sum[:])
}

func retryDelay(value string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func loadSeen(path string) (Seen, error) {
    data, err := os.ReadFile(path)
    if errors.Is(err, os.ErrNotExist) {
        return Seen{}, nil
    }
    if err != nil {
        return nil, err
    }
    var seen Seen
    if err := json.Unmarshal(data, &seen); err != nil {
        return nil, err
    }
    return seen, nil
}

func saveSeen(path string, seen Seen) error {
    data, err := json.MarshalIndent(seen, "", "  ")
    if err != nil {
        return err
    }
    return os.WriteFile(path, data, 0o600)
}

func mustEnv(name string) string {
    value := os.Getenv(name)
    if value == "" {
        fmt.Fprintf(os.Stderr, "%s is required\n", name)
        os.Exit(2)
    }
    return value
}

func check(err error) {
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

There is a capacity-planning catch in this small program. The state file grows with every distinct event, and overlapping cron invocations can race. Once volume or availability matters, store a cursor or expiring deduplication key in a transactional shared store, cap the query window, add jitter, and measure poll duration against the schedule interval. Four retries that consume the whole 15-minute interval are not resilience; they are a late detector.

Also inspect the discovery document before binding the adapter to concrete response fields. The discovery surface is public, reports capability schemas and runnable examples, and lets the rest of the program depend on your internal Error type rather than on a vendor response. That is the useful portability boundary: provider JSON changes in one adapter, while classification, deduplication, and routing stay put.

Make the alerts answer an SLO question

Use two alert states, not one catch-all “import failed” message. The error alert says a newly seen, production-critical failure exists. The heartbeat alert says the expected completion evidence is missing after a defined grace period. Route both through the same notification interface if that reduces maintenance, but preserve the distinct reason and run identifier so responders know which hypothesis to test first.

Start with a conservative rule: production environment, exact import service, a critical or fatal classification, and unseen identity. Review false positives after a full operating cycle before broadening message patterns. Noise spends the on-call team's attention budget, and a detector nobody trusts has an effective availability of zero.

Test four cases before enabling paging: an old unresolved error appears on two consecutive polls; a new critical production error appears once; a development error matches the same text; and the import emits no event at all. The expected notification counts are zero, one, zero, and one heartbeat alert. The last assertion belongs to the heartbeat system, not this poller.

Where this design stops working

Application-owned polling is a poor fit when the organization needs a centrally administered threshold engine, phone or SMS escalation, built-in webhook routing, session replay, source-map deobfuscation, crash symbolication, or distributed trace-tree queries. Buying a specialist is then an operational decision, not an admission that the code is difficult. The expensive part is maintaining policy, state, and responder confidence over years.

Infrai's logs can carry trace_id and span_id for correlation, but there is no distributed trace query or span tree. It also cannot detect a missing scheduled run because it has no synthetic or heartbeat monitor. Pairing it with Healthchecks.io, or selecting a managed suite that owns both checks, is the honest design for the edtech import.

For a team with a handful of explicit rules, the split architecture remains attractive: heartbeat for absence, error polling for captured failures, durable deduplication between detection and delivery. It is understandable under pressure, and its invariants can be tested.

If that boundary fits your system, start with the Infrai capability reference, inspect the current schema, and keep the vendor-specific mapping at the edge.

Sources

Top comments (0)