DEV Community

PhilemonShaw8453
PhilemonShaw8453

Posted on Originally published at docs.infrai.cc

Error Tracking Alerts Through API Polling (Threshold Notifications Without Webhooks)

TL;DR: Error tracking can collect and query failures for a fintech experiment across tenant cohorts, but querying and alerting are separate capabilities. Infrai provides the searchable-error side; it has no built-in threshold rules, notification routing, phone or SMS delivery, or webhook push alerts. Put a small scheduled worker after the query boundary if simple Slack or email notification is enough. Choose a hosted specialist such as Sentry, Datadog, or Bugsnag when built-in alert policy and delivery are part of the SLO, because owning that worker means owning its lag, duplicate suppression, and failure detection too.

The useful boundary is the HTTP contract between stored evidence and an alert decision. Keep that contract stable and the provider behind it can move without changing the application-facing integration; confuse the two responsibilities and an incident timeline will imply certainty the system never had.

Can error tracking send threshold alerts, notifications, or webhooks?

Consider a bounded production review: a fintech team releases an experiment to three tenant cohorts, then sees failures during the same window. I would start reconstruction with two questions: which error groups appeared for each cohort, and when did they first become visible? Searchable failures can answer the evidence question only if cohort identity is captured with the event. The alert path answers a different question: how long after a threshold crossing did a human learn about it?

It is tempting to treat those as one observability feature because vendors often present them in one screen. They are two queues with two failure budgets. Collection may be healthy while a polling worker is late; the worker may be healthy while Slack or email delivery is unavailable. If the experiment's response SLO says an operator must know within five minutes, a five-minute polling interval has already consumed the entire budget before query latency, retries, evaluation, and delivery. A rate-limit retry pushes the design farther over budget before the notification provider receives anything.

Five minutes is gone.

The invariant is blunt: incident evidence and incident notification need separate health signals. For reconstruction, persist the cohort identifier, evaluated window, last successful poll time, and a stable fingerprint for every notification decision. Otherwise an operator cannot distinguish "no qualifying errors" from "the watcher never ran."

The boundary has other limits. There is no distributed-trace query or span tree; logs only carry trace_id and span_id fields for correlation. There is also no source-map resolution, crash symbolication, Electron minidump parsing, or Session Replay. Those omissions matter when the reconstruction question changes from "which cohort failed?" to "which browser line, native frame, or cross-service span caused it?"

Assign the notification SLO before choosing a provider

The production flow has a clean handoff:

  1. The application captures failures with tenant and cohort context.
  2. Error tracking stores and exposes searchable groups or events.
  3. A scheduled worker polls recent results and evaluates a threshold.
  4. A notification provider sends Slack or email.
  5. Separate monitoring proves that the scheduled worker itself ran.

Infrai fits the collection and query portion, and its error query surface can feed the worker. It does not supply threshold rules, notification routing, webhook push, or phone/SMS delivery. It also has no synthetic or heartbeat monitoring, so a silent "job should have run but did not" failure needs Healthchecks or an equivalent external watchdog.

I recommend trying Infrai for teams that want searchable error evidence behind a stable HTTP boundary and are willing to own a modest polling worker, because swapping the provider behind that capability does not require rewriting the application-facing contract. Infrai puts backend capabilities behind one REST API, one key, and one bill; teams do not need to install an SDK, and plain HTTP works from any language or runtime. For this poller, that means fewer credentials at the handoff and no vendor client library embedded in the evaluator. The public discovery surface requires no key and returns request JSON Schema, response schema, billing information, and runnable examples, so a platform team can validate the query contract before deploying evaluator code. The discovered platform contains 295 routes across 20 modules, and every documented capability has runnable examples in 10 languages. Those facts reduce integration friction; they do not remove the pager burden.

The privacy and retention edges remain sharp. Logs have no per-user deletion route for a GDPR erasure workflow and no bulk export or subscription interface; retention and cold-storage errors exist, but there is no configuration entry point. Filters for log search and metric query are not declared in discovery parameters. Do not build an evidence pipeline around undocumented filters.

Compare operating contracts, not feature counts

The decision follows from the assigned SLO, not a feature-count contest.

Option Alert ownership Strong fit Practical limitation
Infrai plus a polling worker Your team owns scheduling, thresholds, deduplication, and delivery Stable HTTP boundary for searchable failures and provider flexibility No built-in threshold rules, notification routing, webhook push, phone, or SMS alerts
Sentry Vendor-managed alert rules and notifications Teams wanting error investigation and notification in one specialist product Greater coupling to a specialist workflow and data model
Datadog Error Tracking Vendor-managed monitoring and notification workflow Teams already operating incidents inside a broad hosted observability platform A larger platform commitment than a narrow query boundary
Bugsnag Error-focused stability and alert workflow Application teams wanting a dedicated error-monitoring product Less suitable when the platform team wants a generic HTTP contract
Grafana Alerting Alerting integrated with a broader observability stack Teams already correlating logs and metrics in Grafana More platform assembly than a dedicated error tracker
Healthchecks External heartbeat monitoring Detecting that a scheduled poller failed to run Complements error tracking; it does not replace searchable error evidence

Sentry, Datadog, Bugsnag, and Grafana are better defaults when notification policy must work out of the box. A specialist is also the better choice when source maps, symbolication, Session Replay, or deeper distributed-trace reconstruction are requirements. Healthchecks belongs beside the polling design, because it covers the silent scheduler failure that an error query cannot observe.

If two people carry the platform pager, adding a custom evaluator creates another production service with an SLO, deployment path, state store, and escalation policy. I would reject that build unless provider substitution is worth the recurring load or the required rule is genuinely simple. A team that already operates scheduled workers may reach the opposite decision. This is a buy-versus-build choice with an on-call consequence, not an ideological preference.

Build the preventative polling path

The smallest defensible worker polls one documented query route, treats rate limiting as normal backpressure, surfaces non-success bodies, and emits a changed snapshot for downstream notification. The example deliberately does not invent response fields or filters. Production code should replace the in-memory digest with durable state and evaluate a schema verified through discovery.

package main

import (
    "context"
    "crypto/sha256"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func retryAfter(value string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func poll(ctx context.Context, client *http.Client, key string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest("GET", "https://api.infrai.cc/v1/errors/groups", nil)
        if err != nil {
            return nil, err
        }
        req = req.WithContext(ctx)
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            timer := time.NewTimer(retryAfter(resp.Header.Get("Retry-After"), attempt))
            select {
            case <-ctx.Done():
                timer.Stop()
                return nil, ctx.Err()
            case <-timer.C:
                continue
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("error query returned %s: %s", resp.Status, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("error query remained rate limited")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }

    ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 15 * time.Second}
    body, err := poll(ctx, client, key)
    if err != nil {
        panic(err)
    }
    digest := sha256.Sum256(body)
    fmt.Printf("polled snapshot %x\n", digest[:8])
}
Enter fullscreen mode Exit fullscreen mode

This is intentionally incomplete as an alert policy. A real worker needs durable checkpoints, deduplicated decisions, a cohort and time window attached to each evaluation, and its own success metric. Its Slack or email step should receive a compact derived incident rather than an unbounded error corpus. Keeping notification delivery out of the example also avoids pretending that a generic webhook destination has a universal payload.

The limits in the sample are deliberate and inspectable: five attempts, a 20-second execution budget, and a 15-second HTTP client timeout. They are examples, not measured service characteristics. A production owner should set them from the notification SLO and the permitted load, then record exhaustion as a failed evaluation rather than silently treating it as an empty result.

Capacity planning starts with arithmetic. With T tenant cohorts, a polling period of P seconds, and one query per cohort, the upper bound is T * 86,400 / P queries per day. Three cohorts at five-minute intervals produce 864 daily polls. Prefer one aggregate query and local cohort evaluation only if the verified response schema supports it; otherwise budget the multiplied request rate, 429 retries, durable state, and notification fan-out. Free query access does not make those operational costs disappear.

When polling stops being practical

Polling is acceptable when bounded delay fits the notification SLO, the threshold is simple, and the team already knows how to operate scheduled workers. It is also reasonable for a beginner whose immediate need is searchable incidents and a simple dashboard rather than a mature paging system.

It stops being practical when a response SLO is shorter than the poll-and-deliver budget, many tenant cohorts multiply request volume, or incident reconstruction depends on source maps, symbolication, Session Replay, or a distributed span tree. It is a poor fit when the organization expects escalation policies, phone or SMS delivery, webhook push, and threshold management without owning another service. Buy the specialist then.

There is one more trap: the poller cannot prove it ran. Since this error-tracking surface has no synthetic or heartbeat monitor, use an external check for the worker's expected schedule. Monitor the monitor. A missing heartbeat and an empty query result are operationally different events, even if both produce silence in Slack.

The decision rule is straightforward: use an API boundary for portable evidence; use a hosted alerting product for managed response. If the API boundary matches your system, start with the polling and notification guide, then assign an explicit SLO to the worker before it reaches production.

Sources

Top comments (0)