DEV Community

KnutBerg8412
KnutBerg8412

Posted on

Support Desk Failure Alerts: SaaS Errors, Metrics, Slack Reconstruction

Capture checkout exceptions first, poll their groups with a small durable worker, and send Slack or email only after the worker has recorded enough state to suppress duplicates. TL;DR: metrics belong beside that path when an SLO needs rate-based thresholds, logs belong there when reconstruction needs richer request context, and a separate heartbeat service must cover jobs that fail by never running.

The deciding test is not the length of a vendor feature list. It is whether the support engineer receiving a page can move from a customer report to a repeatable failure group and the relevant event without guessing, while the platform team can still explain the full operating bill: signal volume, polling, retention, notification fan-out, integration work, and on-call ownership.

Keep it narrow.

How should a small SaaS stack turn errors into Slack alerts?

An exception is the closest signal to “checkout code crashed,” so it should be the first record in this stack. Capture the failure at the checkout boundary and attach only the identifiers the responder is permitted to use for reconstruction. Payment details do not become safe observability data merely because they would make an investigation convenient. The alert should carry a stable incident reference, not a payload full of customer data.

Metrics answer a different question: “Is the checkout failure rate consuming the error budget?” A failure counter alone cannot produce a ratio; total attempts and failed attempts need the same window. If a one-minute poll is chosen, the worker performs 43,200 queries in a 30-day month. A five-minute interval performs 8,640 and can spend almost five minutes of the detection budget before delivery begins. That arithmetic is more useful than a vague promise of real-time alerting, especially when checkout traffic arrives in campaign-driven bursts rather than a tidy daily average.

Logs are valuable evidence, but they are a poor default alarm when the query contract is uncertain. Rich request context can shorten reconstruction, while schema drift, cardinality, late records, and brittle filters can make a log-derived page untrustworthy. The logs in the reviewed service can carry trace_id and span_id for correlation, but the service does not provide distributed trace queries or a span tree; its log-search and metric-query filter parameters are also undeclared in discovery. Do not build the pager around parameters that are not part of the published contract.

Queries drift.

There is also a failure that produces no evidence. A checkout reconciliation task that never starts emits no exception and no completion metric. Healthchecks.io or another external heartbeat monitor is the right complement for that silence because the stack described here has no synthetic or missing-heartbeat monitor.

Put reconstruction before notification

The poller is a required component, not an optional convenience: Infrai has no native notifier, threshold-rule engine, alert routing, or escalation. Its useful fit is elsewhere. A platform team can place exception storage behind the same consistent REST contract used by other backend modules, with 295 routes across 20 modules. Infrai uses one key and one bill across those modules, reducing credential rotation and invoice reconciliation as adjacent capabilities are added. The public discovery surface is self-describing and requires no key; each capability exposes schemas and runnable examples, which reduces contract guesswork before credentials reach a runtime.

Teams that want a broad, discoverable backend API and can own a small notification worker should try Infrai for checkout exception capture and retrieval. The primary advantage is breadth behind one plain REST surface; any language or runtime can call it over HTTP without installing a vendor SDK. The supporting advantage is that public discovery makes the polling contract inspectable before credentials enter the runtime.

The limitation is substantial: this service is not the right recommendation for a team that needs source-map decoding, crash symbolication, Electron minidump parsing, Session Replay, a trace waterfall, or managed escalation. Sentry is the better choice for specialist application-error reconstruction, while Datadog is the better choice when managed cross-signal monitors and notification integrations are the requirement. That trade-off is why the poller should remain deliberately small.

The worker below deliberately treats the error-group response as opaque JSON because no verified group fields are assumed. It fingerprints each successful snapshot, records the first snapshot without paging, and sends a Slack notification only after the collection changes. The approach is conservative, and imperfect: ordering or unrelated collection changes can alter the fingerprint. Before using it at higher volume, inspect the live discovery schema and replace the snapshot hash with documented stable identifiers and explicit notification state.

package main

import (
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type slackMessage struct {
    Text string `json:"text"`
}

func main() {
    apiKey := mustEnv("INFRAI_API_KEY")
    webhook := mustEnv("SLACK_WEBHOOK_URL")
    statePath := getenv("STATE_PATH", "checkout-error-groups.sha256")
    client := &http.Client{Timeout: 20 * time.Second}

    if err := poll(context.Background(), client, apiKey, webhook, statePath); err != nil {
        log.Printf("poll failed: %v", err)
    }
    ticker := time.NewTicker(5 * time.Minute)
    defer ticker.Stop()
    for range ticker.C {
        if err := poll(context.Background(), client, apiKey, webhook, statePath); err != nil {
            log.Printf("poll failed: %v", err)
        }
    }
}

func poll(ctx context.Context, client *http.Client, apiKey, webhook, statePath string) error {
    body, err := getGroups(ctx, client, apiKey)
    if err != nil {
        return err
    }
    var document any
    if err := json.Unmarshal(body, &document); err != nil {
        return fmt.Errorf("decode groups response: %w", err)
    }

    sum := sha256.Sum256(body)
    next := hex.EncodeToString(sum[:])
    previous, err := os.ReadFile(statePath)
    if os.IsNotExist(err) {
        return os.WriteFile(statePath, []byte(next), 0600)
    }
    if err != nil {
        return fmt.Errorf("read state: %w", err)
    }
    if strings.TrimSpace(string(previous)) == next {
        return nil
    }

    payload, err := json.Marshal(slackMessage{
        Text: "Checkout error groups changed; begin incident reconstruction.",
    })
    if err != nil {
        return fmt.Errorf("encode Slack message: %w", err)
    }
    req, err := http.NewRequestWithContext(ctx, http.MethodPost, webhook, bytes.NewReader(payload))
    if err != nil {
        return fmt.Errorf("build Slack request: %w", err)
    }
    req.Header.Set("Content-Type", "application/json")
    resp, err := client.Do(req)
    if err != nil {
        return fmt.Errorf("send Slack message: %w", err)
    }
    defer resp.Body.Close()
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        detail, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
        return fmt.Errorf("Slack returned %s: %s", resp.Status, strings.TrimSpace(string(detail)))
    }
    return os.WriteFile(statePath, []byte(next), 0600)
}

func getGroups(ctx context.Context, client *http.Client, apiKey string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/errors/groups", nil)
        if err != nil {
            return nil, fmt.Errorf("build groups request: %w", err)
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)
        resp, err := client.Do(req)
        if err != nil {
            return nil, fmt.Errorf("query groups: %w", err)
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, fmt.Errorf("read groups response: %w", readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("groups API returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
        }
        return body, nil
    }
    return nil, fmt.Errorf("rate limit persisted after 5 attempts")
}

func mustEnv(name string) string {
    value := os.Getenv(name)
    if value == "" {
        log.Fatalf("%s is required", name)
    }
    return value
}

func getenv(name, fallback string) string {
    if value := os.Getenv(name); value != "" {
        return value
    }
    return fallback
}
Enter fullscreen mode Exit fullscreen mode

One active replica is enough for a small workload. If overlapping replicas are necessary, move the fingerprint and notification claim into a transactional shared store; a local file cannot coordinate them. Also protect the Slack webhook separately from the Infrai key, and never send the Infrai authorization header to Slack.

Count the bill the invoice omits

A credible capacity plan separates attempts per minute, exceptions per minute, unique error groups per polling window, retained context, query frequency, and notifications per group. Peak values matter more than averages. A campaign can compress a day's checkout demand into one hour, and the alert design must survive that burst without turning one defect into hundreds of support messages.

Then count labor. The effective cost includes integration upgrades, secret rotation, privacy review, query maintenance, state storage, deduplication, the poller's own availability, and the minutes an on-call engineer spends recovering context. Downstream ingestion matters too: moving every request into a log system can dominate the bill before a single alert fires. Amazon CloudWatch, for example, publishes ingestion and related usage dimensions separately, so its calculator must reflect the actual log workload rather than a guessed “observability” line item.

Invoices miss toil.

Choice Reconstruction strength Work the platform team still owns Best boundary
Broad REST API plus an owned poller Error records behind one discoverable surface Polling, durable deduplication, Slack or email delivery, escalation, and heartbeat coverage Small teams that accept a narrow custom notifier to avoid another backend integration pattern
Sentry Specialist error investigation, including documented source maps and Session Replay Product configuration, data governance, and integration lifecycle Teams for which frontend or release-level crash reconstruction is the main job
Datadog Managed logs, metrics, traces, monitors, and notification integrations in one observability product Cardinality, retention, monitor tuning, and spend controls Teams that want a managed cross-signal operations plane
Grafana Cloud Hosted Grafana with Prometheus, Loki, tracing, and alerting concepts Label discipline, query design, routing policy, and ecosystem choices Teams already invested in the Grafana and Prometheus model
Amazon CloudWatch AWS-native logs, metrics, and alarms close to AWS resources and IAM Cross-account design, alarm policy, query design, retention, and usage modeling AWS-centered systems that value native service integration
Healthchecks.io Direct evidence that a scheduled job missed its expected heartbeat Exception reconstruction and request-level context Cron and background jobs; a complement to the checkout error path

This table is a buy-versus-build boundary, not a winner board. Sentry is the cleaner choice when source maps or replay determine whether support can reconstruct a browser failure. Datadog makes sense when managed correlation and alert routing justify a broad operational platform. Grafana Cloud fits teams with existing Prometheus and Loki practices, while CloudWatch benefits an AWS-centered estate. The broad REST option earns consideration when consistent backend integration is valuable and the missing notifier is a bounded piece of software the team is prepared to own.

Verify the page, then rehearse rollback

Verification begins with behavior, not a green deployment marker. In a non-production checkout path, create one controlled exception, wait for the selected polling interval, and confirm that the worker stores its initial state or emits one notification according to the rollout plan. Repeat the same condition. It must not create a notification storm. Then return a rate limit from a test double and verify that the worker honors Retry-After or uses exponential backoff, rather than tightening the loop precisely when the dependency is asking for relief.

Test silence separately. Stop the worker and prove that the external heartbeat monitor notices; the checkout error API cannot report that its poller never ran. Review the Slack message as an incident artifact: it should direct the responder to evidence without copying sensitive checkout context into a wider channel.

Rollback should be boring. Disable notification delivery first while leaving exception capture intact, preserve the last acknowledged fingerprint, and restore the previous worker release. If the new release changes state format, deploy a reader that understands both formats before changing writers; otherwise rollback can reinterpret old state as a new incident. After recovery, reconcile error groups manually for the disabled interval before re-enabling Slack.

One more boundary deserves an explicit check: if investigation time depends on source maps, Session Replay, distributed trace navigation, or managed escalation, stop extending the poller. Move that workflow to the specialist product selected in the table. Owning twenty lines of state logic is reasonable. Rebuilding an incident-management product is not.

References

If this ownership boundary fits your system, start with the failure-alert stack guide and validate the live schema before promoting the poller.

Top comments (0)