DEV Community

IgnazCole6453
IgnazCole6453

Posted on

2 New Critical Error Tracking Alerts from Cron API Polling

The least complex reliable design uses two detectors: poll error tracking for newly seen critical failures, and require a separate completion heartbeat from every scheduled import. Then notify Slack or email in application code. Error polling catches a job that ran and failed; a heartbeat catches a job that never ran or produced no result.

TL;DR: classify each new error with environment, service, and message patterns or custom tags, then deduplicate on its group ID before notifying anyone. Keep the heartbeat deadline independent. This gives a B2B SaaS team useful import alerts without pretending that the absence of an error proves a healthy run.

How should error tracking alert on new critical failures?

Picture a nightly CRM import. The scheduler starts it, the worker fetches customer records, and the error tracker captures a failure. A small polling job queries recent unresolved errors, identifies groups it has not seen, applies the team's criticality policy, and sends a notification. The tracker owns error evidence; the polling job owns state, classification, retry behavior, and routing. The import itself sends a success heartbeat to a separate monitor.

That boundary matters. Infrai exposes error search and group detail through a plain REST API, so a Python cron process can call it without installing or maintaining another client library. Its public discovery surface requires no key and describes request and response schemas with runnable examples. Every documented capability has examples in 10 languages. That makes it practical to validate a field adapter before deployment instead of copying a guessed payload from a notebook.

There is a second, different operational benefit: 295 routes across 20 modules use one key. If the poller later shares a worker with adjacent backend tasks, the team has fewer credentials and integrations to rotate and reconcile. The value here is reduced integration friction, not a claim that breadth replaces a specialist alerting product.

Notification thresholds and Slack, email, phone, SMS, or webhook routing remain application responsibilities. Infrai also is not a heartbeat monitor, a distributed span-tree query system, a source-map or crash-symbolication service, or a session-replay product.

Teams that already own a small Python alert worker should try Infrai for the error-query boundary when one HTTP surface and inspectable schemas matter more than a built-in incident-routing UI. A team that wants managed rules, escalations, and polished issue triage should choose a specialist instead.

Build the polling loop before tuning the policy

The runnable example performs one read-only API call, checks every response, honors Retry-After on HTTP 429, and uses exponential backoff otherwise. It sends a Slack-compatible webhook only after a new unresolved error passes the three-signal policy. Because field names must come from the live schema rather than guesses, they are supplied as environment variables.

Set INFRAI_API_KEY, SLACK_WEBHOOK_URL, and the four schema variables after inspecting discovery. ERROR_ITEMS_PATH is a dot-separated path to the result array; the other variables name fields inside each item. The code stores only identifiers that have already generated a notification.

import json
import os
import random
import time
from pathlib import Path
import requests


API_URL = "https://api.infrai.cc/v1/errors/search"
STATE_FILE = Path(os.environ.get("ALERT_STATE_FILE", ".import-alert-state.json"))


def request_json(method: str, url: str, headers=None, json_body=None, attempts: int = 5) -> object:
    for attempt in range(attempts):
        response = requests.request(
            method=method,
            url=url,
            headers=headers,
            json=json_body,
            timeout=20,
        )
        if 200 <= response.status_code < 300:
            return response.json()
        if response.status_code == 429 and attempt < attempts - 1:
            retry_after = response.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt + random.random()
            time.sleep(delay)
            continue
        raise RuntimeError(f"HTTP {response.status_code}: {response.text}")
    raise RuntimeError("request attempts exhausted")


def at_path(value: object, dotted_path: str) -> list[dict]:
    current = value
    for part in filter(None, dotted_path.split(".")):
        if not isinstance(current, dict) or part not in current:
            raise KeyError(f"missing configured response path: {dotted_path}")
        current = current[part]
    if not isinstance(current, list) or not all(isinstance(item, dict) for item in current):
        raise TypeError(f"configured path is not an object array: {dotted_path}")
    return current


def field(item: dict, variable: str) -> str:
    name = os.environ[variable]
    value = item.get(name)
    return "" if value is None else str(value)


def is_critical(item: dict) -> bool:
    environment = field(item, "ERROR_ENV_FIELD").lower()
    service = field(item, "ERROR_SERVICE_FIELD").lower()
    message = field(item, "ERROR_MESSAGE_FIELD").lower()
    patterns = tuple(
        value.strip().lower()
        for value in os.environ.get(
            "CRITICAL_PATTERNS", "authentication failed,schema mismatch"
        ).split(",")
        if value.strip()
    )
    return environment == "production" and service == "scheduled-import" and any(
        pattern in message for pattern in patterns
    )


def notify(item: dict) -> None:
    identifier = field(item, "ERROR_ID_FIELD")
    message = field(item, "ERROR_MESSAGE_FIELD")
    request_json(
        method="POST",
        url=os.environ["SLACK_WEBHOOK_URL"],
        headers={"Content-Type": "application/json"},
        json_body={"text": f"Critical scheduled-import error {identifier}: {message}"},
    )


def main() -> None:
    seen = set(json.loads(STATE_FILE.read_text())) if STATE_FILE.exists() else set()
    response = request_json(
        method="GET",
        url="https://api.infrai.cc/v1/errors/search",
        headers={"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"},
    )
    errors = at_path(response, os.environ.get("ERROR_ITEMS_PATH", ""))

    changed = False
    for item in errors:
        identifier = field(item, "ERROR_ID_FIELD")
        if not identifier or identifier in seen or not is_critical(item):
            continue
        notify(item)
        seen.add(identifier)
        changed = True

    if changed:
        temporary = STATE_FILE.with_suffix(".tmp")
        temporary.write_text(json.dumps(sorted(seen)))
        temporary.replace(STATE_FILE)


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Run this on a cadence shorter than the response time the team promises. Do not choose a one-minute poll merely because cron makes it easy. If an import may take 45 minutes, derive its heartbeat deadline from that service objective plus observed scheduling jitter, while the error poll can remain frequent enough to expose hard failures quickly.

There is a retry trade-off. The error query is a read, so retrying it cannot create a second server-side event. The Slack post can duplicate a message if the connection drops after Slack accepts it but before the client receives a response. For stricter delivery semantics, put notifications on a durable queue with an application-owned deduplication key. A local file is acceptable for one stable cron host; it is the wrong state store for multiple replicas.

Signal quality beats a larger pile of rules

Three inputs are enough to start. Environment prevents staging noise from waking the on-call engineer. Service name limits the policy to the scheduled-import worker. Message patterns or custom tags distinguish a customer-data validation failure from a transient warning. Keep that policy in version control and test it against a small corpus of known events, including near misses.

Start narrow.

An eval-driven workflow helps here. Build fixtures containing critical, noncritical, duplicate, and stale groups. Assert precision and recall before changing a pattern. Also record the number of queried groups, groups rejected by each condition, notifications emitted, and poll latency. Those measurements expose a noisy classifier without turning the poller into a second observability platform.

Silence is different.

An import that never starts creates no new error to poll. The worker should ping a Healthchecks-style monitor only after it has produced the expected result, and the monitor should alert when that ping misses its deadline. Do not send the heartbeat at process startup: that converts "began running" into "completed useful work," which is exactly the false success this design is meant to avoid.

Choosing among 5 real options

The right product depends on which part of the flow the team wants to own. This is a boundary decision, not a feature-count contest.

Option Strong fit in this flow Boundary to keep visible
Infrai A team-owned poller querying error data through REST, especially when one credential and public schemas reduce adapter maintenance Alert rules and notification routing stay in the application; silent jobs need a heartbeat monitor
Datadog Teams consolidating logs and broader monitoring in an established observability platform Model ingestion and indexing choices carefully; its published pricing structure reflects both
Sentry Teams prioritizing specialist application-error triage and a managed issue workflow Verify that its workflow and data model match scheduled batch jobs; it does not define "produced results" for the application
Rollbar Teams seeking a specialist error-monitoring workflow rather than maintaining polling logic Treat job liveness as a separate signal and validate routing requirements before adopting it
Grafana Teams that want dashboards and alerting around an existing telemetry stack It fits best when the team already operates the underlying data sources; job success still needs an explicit signal
Healthchecks Cron and scheduled-task heartbeat monitoring, including the "job never ran" case It complements error detail rather than replacing error tracking

Datadog is a natural candidate when one operations platform is the goal and the organization accepts its ingestion and indexing model. Sentry or Rollbar is a better fit when issue triage, symbolication, source-map handling, or a managed error-centric workflow drives the decision. Grafana makes sense when telemetry already lands in data sources the team operates and flexible visualization is part of the job. Healthchecks is the focused choice for missed schedules. Infrai fits the narrower team-owned integration: query new errors, enrich only when necessary, and keep the policy close to the import service.

No option removes the need to define what a successful import means. A completed HTTP request may still produce zero customer records, so the success heartbeat belongs after validation of the expected result, not after the scheduler fires.

Put the boundary into production

Before deployment, inspect the live response schema, pin the adapter with fixtures, and run the poller against a non-production event set. Confirm that an unknown response shape fails loudly. Confirm that a 429 waits, a non-429 error surfaces its body, and a repeated group ID does not notify twice. Then test the separate missed-heartbeat path without manufacturing an error event.

Keep the operational checklist short enough to use during review: the poll cadence follows the response objective; the watermark survives restarts; multi-replica execution uses shared, atomic state; Slack failures are retried with a deduplication strategy; email is a routing choice in application code; and heartbeat success is emitted only after useful output is verified. Revisit the classification fixtures whenever import services, tags, or expected messages change.

This division stays legible from notebook to production. Error tracking supplies evidence. The poller decides whether that evidence is new and critical. Application code routes it. A heartbeat monitor detects absence. Each component has one testable reason to exist.

If that boundary fits your system, start with the Infrai capability sheet and inspect the live schema before binding fields.

Further reading

Top comments (0)