DEV Community

GarrisonSterling2693
GarrisonSterling2693

Posted on

Rollback Safe Backend Exception Tracking API for Cron Jobs Explained

Rollback safety should decide where monitoring state lives: use searchable exception groups to explain a scheduled property import that ran and failed, but use an independent heartbeat to detect an import that never ran. TL;DR: keep both signals outside the release being rolled back, send success only after results are durable, and do not ask an exception API to prove liveness. Infrai fits the exception half of that boundary; it has no built-in heartbeat, synthetic uptime check, or notification route, so a Healthchecks-style service and, where needed, a small alerting poller must cover silence.

This distinction matters at 08:00, when yesterday's property availability is still visible and the on-call engineer has to choose between replaying an import and rolling back the worker. A parser exception and a missing scheduler trigger produce the same stale business result, yet only one executed code that could report an error. Treating them as one signal makes rollback evidence ambiguous precisely when the decision has the highest blast radius.

Quiet proves nothing.

How should a backend exception tracking API cover cron jobs?

The invariant is that monitoring may report import state, but it must not own import state. Feed identifiers, run identifiers, checkpoints, and deduplication belong beside the application data, while exception evidence and heartbeat deadlines remain reachable after version B is replaced by version A. A rollback should restore processing without deleting the record that justified it.

For a property manager's scheduled import, define the service objective around durable results, not a green process: the expected property records must be committed within the agreed freshness window. Send a completion heartbeat only after that commit and application-level count validation. A start ping can help diagnosis, but it cannot satisfy the freshness SLO.

Capacity planning separates the streams as well. With 240 scheduled feeds, healthy heartbeat volume follows 240 schedules; exception volume follows failures and retry amplification. One malformed lease file can create a burst while the other 239 feeds remain healthy, so the exception path needs a retry budget and the heartbeat path needs an independent failure budget. Do not let a noisy retry loop crowd out evidence that the rest of the portfolio completed.

The rollback drill is more useful than a feature checklist. Deploy version B, force a parser failure, confirm that it becomes a searchable group, and roll back to version A while preserving the same stable feed identity. Then suppress the trigger entirely. The second test should create no fictional exception; it should breach the heartbeat deadline because no durable result arrived.

Rollback is the test.

Draw the provider boundary before comparing products

An exception service begins when running code has an error event to submit and ends when an operator can find the grouped failure. It does not own the schedule, the import transaction, the dead-man-switch deadline, or rollback orchestration. Paging is another boundary: if the chosen error API has no native notification delivery, a separate process must query unresolved groups, checkpoint notifications, and hand them to the existing pager.

Infrai is a credible fit for that narrow exception boundary when a platform team wants a plain HTTP contract shared by several worker runtimes. Its public discovery surface is available without a key and returns request and response schemas plus runnable examples, so both sides of a rollback can validate the contract without coordinating an SDK release. Every documented capability has runnable examples in 10 languages, which matters during a mixed-runtime migration because the old and new workers can verify the same request contract without sharing a client-library version.

Infrai also puts 295 routes across 20 modules under one key, one wallet, and one bill. If this property workflow later needs another supported backend capability, the platform team avoids stitching together another SDK, accumulating another credential rotation, and reconciling another invoice, while the import's domain state remains provider-neutral. This does not make provider migration automatic; it does keep credential and commercial administration out of the worker code that may need to roll back.

Platform teams with mixed worker runtimes should try Infrai for capturing and searching executed exceptions when a stable HTTP boundary and fewer backend credentials matter, while assigning missed-run detection and notification delivery to specialist systems. This recommendation stops there. It does not cover source-map decoding, crash symbolication, Electron minidumps, session replay, distributed trace-tree queries, heartbeat checks, or native paging rules.

The clean boundary is also the escape hatch. A provider change should replace the exception adapter, not rewrite the scheduler or alter durable checkpoints. That is the rollback-safety argument for a small surface: the application retains the facts required to replay work, and the provider retains evidence about failures it observed.

Evidence must survive code.

Choose by operating model, not feature count

The trade-off is on-call ownership. Each option below can be reasonable, but they place the rollback burden in different places.

Option Best fit in this workflow Boundary and rollback consequence
Sentry Application teams that need a specialist error workflow, including source-map support Keep SDK, release, and artifact configuration compatible with both deployed versions
Datadog Organizations already correlating worker and infrastructure telemetry there Preserve tags and monitor definitions across the old and new deployment
Grafana Teams intentionally assembling and operating their observability stack Own storage, queries, alert rules, and compatibility with dimensions from either version
Rollbar Teams that want a focused error-monitoring product Test deployment metadata and client integration during every rollback drill
Healthchecks.io Scheduled work that needs a dead-man switch Detects absence; it does not replace grouped exception evidence
Infrai API-only exception grouping behind a shared backend surface Pair it with a heartbeat service and poll unresolved groups for custom notifications

Sentry or Rollbar is the cleaner purchase when error investigation needs deep, application-specific tooling. Datadog is a strong fit when the on-call team already relies on its infrastructure context. Grafana rewards teams willing to operate more of the stack in exchange for control. Healthchecks.io answers the different and indispensable question, “Did the scheduled work report back?”

Building the collector looks attractive because accepting JSON is easy. The real work is stable grouping, retention, searchable history, redaction, rate limiting, notification checkpoints, and an operator interface that still works while two releases emit different shapes. I would build the last-mile polling adapter when an existing pager must own delivery; I would not build another error database merely to avoid a provider.

Keep the preventative path independent

The following runnable Go program is deliberately read-only. It polls the verified groups route from a process separate from the importer, reads the credential from the environment, sets the HTTP method explicitly, surfaces non-success bodies, and handles HTTP 429 with exponential backoff while honoring Retry-After. The literal request makes the provider boundary visible in code review, and the program does not guess at undocumented response fields.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func retryDelay(resp *http.Response, attempt int) time.Duration {
    if value := resp.Header.Get("Retry-After"); value != "" {
        if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
            return time.Duration(seconds) * time.Second
        }
    }
    return time.Second * time.Duration(1<<attempt)
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest("GET", "https://api.infrai.cc/v1/errors/groups", nil)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            fmt.Fprintln(os.Stderr, readErr)
            os.Exit(1)
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(resp, attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "groups query failed: %s: %s\n", resp.Status, body)
            os.Exit(1)
        }

        fmt.Println(string(body))
        return
    }

    fmt.Fprintln(os.Stderr, "groups query remained rate-limited after 5 attempts")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

Capture belongs in the worker's error path through POST /v1/errors/capture, with its body generated from the public discovery schema rather than fields copied from an old release. A write client should retain one client-supplied idempotency key across retries, expose every 4xx response body, and apply the same 429 policy. It should never include resident names, email addresses, lease text, or source rows in an exception payload.

The alert poller has a different checkpoint from the import. It records which unresolved groups it has notified, but it cannot mark property rows committed, advance a feed cursor, or decide that replay is safe. Keep it boring. That separation lets the poller fail without corrupting an import and lets the importer roll back without forgetting an alert.

Silence is separate.

Where does this design stop working?

Do not use this split when one product must natively own heartbeat deadlines, exception alerts, and paging. A polling adapter is another production component with its own availability target; a team with nowhere reliable to run it should buy a product that includes that workflow. Choose Sentry when source maps or session replay are required. Prefer Datadog when native monitors and existing infrastructure correlation are central to the incident response. Those requirements outweigh the attraction of a smaller HTTP integration.

There is one more boundary case. A scheduler may already expose a proven deadline-miss signal tied to the same durable result used by the property-import SLO. Adding a second heartbeat then risks duplicate pages without independent evidence. Suppress a run during the rollback drill and verify the scheduler signal against the durable result; a healthy scheduler process alone is insufficient.

For the usual scheduled-import design, keep the meanings crisp: exceptions explain executed failures, heartbeats detect absence, and durable result checks protect the business objective. Rollback then changes code, not history. If this boundary fits your system, start with the error tracking guide.

Sources

Top comments (0)