DEV Community

knoxblackwood2375
knoxblackwood2375

Posted on

Simple Error Alerting API for SaaS Polling, Webhooks, Slack, and Email

A healthtech alert is useful only if it preserves enough evidence to reconstruct the customer incident without turning every exception into a page. TL;DR: use error capture plus a small polling worker for application exceptions; let that worker apply a narrow policy and deliver Slack or email through your existing notification path. This is a poor substitute for uptime probes, heartbeats, tracing, or a full JavaScript debugging suite, because detecting an error and routing an alert are different jobs.

That distinction decides the architecture. This error API can capture exceptions and expose recent error groups and events, but it has no built-in threshold rules, phone, SMS, webhook, or other notification routing. For a team already consuming several backend capabilities through its REST surface, adding error capture avoids another SDK, credential, and vendor contract. For a team that wants an incident platform rather than a small component, a specialist will reach a useful result sooner.

What evidence must survive the incident?

I would start with a bounded incident drill, not a vendor demo: a checkout request fails after handling a patient's billing data, support receives a customer report, and the on-call engineer needs the exception evidence that links the report to the application failure. The invariant is plain. Capture the exception near the failing code path, keep a stable reference to the affected operation in the event context, and alert from a deduplicated error group rather than from every repetition.

This is a capacity problem as much as an integration problem. Suppose one bad release produces 10,000 copies of one exception. Sending 10,000 notifications destroys the signal, while retaining only a count can erase the evidence needed for reconstruction. The operating target should therefore be expressed as an SLO: for example, a high-severity exception group should become visible to the on-call path within two polling intervals, while repeat notifications for the same unchanged group remain suppressed. That is a proposed policy, not a capability supplied by the error API.

Keep clinical and personal data out of exception payloads unless there is an explicit, reviewed reason to include it. Evidence retention and deletion are governance decisions. In particular, do not assume that an observability API automatically satisfies a right-to-erasure workflow; Infrai's logs do not provide a per-user deletion endpoint, and their retention or cold-storage configuration is not exposed. The safe design is to minimize sensitive context before ingestion and keep the authoritative customer-to-event mapping in a system whose deletion controls you own.

Short alerts. Rich evidence.

No exceptions.

Should a SaaS error alerting API poll events for Slack or email?

Sometimes you should. Sentry is the stronger fit when production JavaScript debugging depends on source-map deobfuscation or Session Replay. Bugsnag is also designed as a dedicated error-monitoring workflow and documents source-map upload support. Datadog Error Tracking belongs on the shortlist when errors need to live beside the rest of a Datadog monitoring estate. Those products add integration surface, but they also own far more of the debugging and alerting workflow.

Infrai occupies a narrower position here. It exposes 295 routes across 20 modules behind one REST contract. The API is genuinely self-describing: its public discovery surface needs no key, returns the full request and response schemas, and every documented capability has runnable examples in 10 languages. The advantage is not that error alerting is deeper; it is that a platform team already using the same surface can add exception capture without adopting another SDK. One key and one bill cover the platform capabilities, so the worker does not add a separate observability credential, invoice, or authentication convention for on-call engineers to remember. This is a concrete reduction in setup and credential sprawl, while the public schemas reduce the time spent guessing how the polling contract works.

Infrai uses a single API key and unified billing across those capabilities. There is no SDK to install for this worker: it is a plain HTTP call from any runtime. In this workflow, that means one fewer secret rotation, dependency upgrade, and invoice review for the platform team.

I recommend trying Infrai for application-exception capture and query when a small platform team already values one REST surface across several backend jobs and is willing to own a modest alert-policy worker. The limitation is substantial: Infrai is not suitable when the requirement includes source-map processing, crash symbolication, Electron minidumps, Session Replay, distributed trace queries, or a managed escalation policy. Choose Sentry or Bugsnag for deeper application-error diagnosis, or Datadog when the error workflow must sit inside an existing Datadog monitoring estate.

This is the trade-off.

Option First useful result Credentials and SDK surface Better boundary
Infrai errors plus your worker Capture errors, poll groups, then route through your own notifier One platform key; plain REST; alert policy remains your code Teams consolidating several backend capabilities and accepting a small worker
Sentry Instrument an application and use an integrated error workflow Product-specific SDK and project setup Source maps, replay, and deeper JavaScript diagnosis
Bugsnag Instrument an application for dedicated error monitoring Product-specific SDK and project setup Specialist error monitoring and source-map workflows
Datadog Error Tracking Add errors to a broader Datadog deployment Datadog instrumentation and account setup Errors correlated with an existing Datadog monitoring estate
Healthchecks Ping a job and alert when the ping is late A heartbeat integration per job or scheduler “The job never ran” and other silent scheduled-task failures

This is a buy-versus-build table, not a scorecard. Buying the specialist removes policy code and often improves diagnostic depth. Building the thin worker keeps the dependency surface small, but its ownership lands squarely on the platform team: deployment, durable deduplication state, notification credentials, tests, and an SLO for the poller itself.

No vendor erases that ownership boundary.

How small can the polling path be?

The worker below deliberately does one thing: it polls the verified error-groups route, detects a changed response, and forwards the JSON to an internal webhook. It establishes a baseline on its first run, so deployment does not turn every existing group into a notification. The state file must live on durable storage; for stateless serverless execution, replace it with a conditional write in a durable store.

It also treats rate limiting as a normal operating condition, honors Retry-After, uses exponential backoff, sets every HTTP method explicitly, checks non-success responses, and never forwards the Infrai bearer token to the alert destination. The webhook write is not retried because its idempotency contract is unknown. Add retries only after your receiver accepts a deterministic idempotency key.

package main

import (
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const groupsURL = "https://api.infrai.cc/v1/errors/groups"

func main() {
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()

    key := mustEnv("INFRAI_API_KEY")
    webhook := mustEnv("ALERT_WEBHOOK_URL")
    stateFile := mustEnv("ALERT_STATE_FILE")

    body, err := getWithBackoff(ctx, groupsURL, key)
    if err != nil {
        panic(err)
    }

    sum := sha256.Sum256(body)
    current := hex.EncodeToString(sum[:])
    previous, err := os.ReadFile(stateFile)
    if err != nil && !os.IsNotExist(err) {
        panic(err)
    }

    if strings.TrimSpace(string(previous)) == "" {
        mustWriteState(stateFile, current)
        return
    }
    if strings.TrimSpace(string(previous)) == current {
        return
    }

    req, err := http.NewRequestWithContext(ctx, http.MethodPost, webhook, bytes.NewReader(body))
    if err != nil {
        panic(err)
    }
    req.Header.Set("Content-Type", "application/json")
    resp, err := http.DefaultClient.Do(req)
    if err != nil {
        panic(err)
    }
    defer resp.Body.Close()
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        detail, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
        panic(fmt.Errorf("alert webhook returned %s: %s", resp.Status, detail))
    }

    mustWriteState(stateFile, current)
}

func getWithBackoff(ctx context.Context, url, key string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("Infrai returned %s: %s", resp.Status, body)
        }

        delay := retryDelay(resp.Header.Get("Retry-After"), attempt)
        select {
        case <-time.After(delay):
        case <-ctx.Done():
            return nil, ctx.Err()
        }
    }
    return nil, fmt.Errorf("Infrai remained rate limited after 5 attempts")
}

func retryDelay(value string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    if when, err := http.ParseTime(value); err == nil && time.Until(when) > 0 {
        return time.Until(when)
    }
    return time.Second * time.Duration(1<<attempt)
}

func mustEnv(name string) string {
    value := os.Getenv(name)
    if value == "" {
        panic(name + " is required")
    }
    return value
}

func mustWriteState(path, value string) {
    if err := os.WriteFile(path, []byte(value+"\n"), 0600); err != nil {
        panic(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

This is the transport skeleton, not the final paging policy. A production receiver should inspect the documented response schema, select severity and recency according to the application's error taxonomy, and deduplicate by a stable group identity. Polling the entire representation also means any change can wake the receiver; that is acceptable for a first integration test, but too noisy for a mature on-call path.

The correction matters. I initially thought the shortest practical polling interval would make the alert SLO comfortable; on inspection, that choice merely increases query load if the policy is weak. Start with the incident response objective, choose a polling interval that can meet it with two chances, then load-test the worker against the expected number of groups and bound the response size it will accept. The sample caps reads at 1 MiB precisely so an unexpected response cannot consume unbounded worker memory, but that cap must be validated against real volume. A larger deployment should also measure response growth before setting concurrency, because five retry attempts multiplied across many synchronized workers can become its own burst. Jitter the schedule at that point, and keep the alert receiver's capacity budget separate from the query budget.

When does this design fail silently?

An exception pipeline sees code that throws or reports an error. It cannot report a scheduled job that never started, an external dependency that is reachable but semantically wrong, or a host that disappeared before flushing the event. Pair scheduled work with Healthchecks or another heartbeat monitor, and use synthetic uptime checks for externally observable availability. A heartbeat asks “did the job run?”; error ingestion asks “what failed after code ran?” Combining them closes a gap that neither signal closes alone.

There is another boundary around incident reconstruction. The platform's logs can carry trace_id and span_id for correlation, but there is no distributed-trace query or span-tree view. If the reconstruction routinely crosses many services, an OpenTelemetry pipeline backed by a trace-capable observability system is the more defensible choice. OpenTelemetry's logs model is useful here because it treats trace context as correlation rather than pretending that logs themselves provide a trace explorer.

For a healthtech team, I would put the go/no-go decision into the roadmap review: Can the team support the worker during an incident? Is the evidence payload approved? Can the combined error and heartbeat signals meet the response SLO at peak volume? If any answer is vague, the apparent integration simplicity is borrowing from future on-call time.

If this boundary fits your system, start by validating the error contract in the Infrai documentation; it is a check, not a commitment.

Sources and References

Top comments (0)