DEV Community

Hwpgsd503817
Hwpgsd503817

Posted on

MVP Error Tracking: 4 Decisions for Grouping, Search, and Alerting

The page says a Node.js SaaS catalog import has produced zero updates for 45 minutes. The on-call opens the alert and needs three things immediately: whether the Sentry alternative's error-tracking API received a failure, which shop and scheduled run stopped, and whether the run failed or never started.

TL;DR: a simple error-tracking API is a reasonable Sentry alternative for one small SaaS app when the job is exception capture and triage: capture an error, inspect its grouped issue and event payload, search the error set, then resolve the group after the fix ships. It is not a substitute for Sentry-level notification routing, frontend debugging ergonomics, distributed trace exploration, source-map processing, crash symbolication, or Session Replay. More importantly for scheduled imports, an error tracker cannot report a run that emitted no error because it never ran. Pair it with a heartbeat monitor such as Healthchecks, or instrument an explicit result signal and operate the poller yourself.

That separation is the decision. Everything else is procurement detail.

What should have fired before the zero-results page?

Work backward from the page. A customer-facing query or a downstream freshness check notices zero catalog changes, but by then the detection delay includes the schedule interval, the job runtime, and whatever lag exists before someone looks at the result. The earlier signal is a per-run outcome with a stable run ID, tenant ID, scheduled time, completion time, result count, and status. If the worker throws, error tracking captures the exception and preserves the event payload for triage. If no worker starts, only a missing heartbeat exposes the silence.

There are therefore two failure modes, and combining them into an error_count > 0 alert creates a blind spot:

  1. Loud failure: the import starts and raises an exception. Grouping reduces repeated events into an issue, while event detail carries the run and tenant context.
  2. Silent failure: the scheduler, queue handoff, or worker never produces an outcome. A deadline-based heartbeat must alert on the missing completion.

The SLO should describe the user-visible outcome, not the tool: for example, “scheduled imports produce a terminal outcome within the agreed completion window.” The exact window belongs in the service contract and capacity model; it cannot be inferred from an error-tracking product. Paging at 45 minutes is sensible only if the normal and high-percentile runtimes leave enough response budget before the SLO is breached.

This is where a tidy error dashboard can mislead an infrastructure lead. Zero new exceptions may mean zero failures. It may also mean zero execution.

Instrument the outcome, not merely the exception

For an e-commerce importer, attach the same correlation fields to the schedule dispatch, worker completion, and captured exception. tenant_id, import_id, and run_id make cost and ownership attributable; trace_id and span_id can correlate records where available, but fields alone do not create a distributed trace query experience or a span tree.

The result signal should be emitted once at the terminal boundary, after the importer knows the number of accepted records. Do not manufacture an exception for a legitimate zero-result run. Record the count, then let the business rule decide whether zero is expected for that tenant and feed. Otherwise a low-volume merchant turns into a recurring false page.

Capacity planning matters here. A poller that searches every tenant independently multiplies query traffic with tenant count, while one batched control-plane check reduces calls but increases the blast radius of a broken poll cycle. Whichever shape you choose, expose the poller's own successful evaluation time. An alerting process with no self-signal merely relocates the silent-failure problem.

Sensitive payloads need a boundary too. Error events are diagnostic records, not a second customer database. Avoid credentials, payment data, and raw personal data; the absence of a per-user log-deletion interface makes a casual “send everything” policy particularly hard to reconcile with deletion obligations. OWASP's logging guidance is the useful floor for deciding what must be excluded or masked.

Should a Node.js SaaS use Sentry or a simple error-tracking API?

The honest comparison is not “which product can receive an exception?” All of them can occupy some part of that workflow. The useful questions are who owns notification delivery, how much frontend context is required, and whether a missing run is a first-class event.

Option Best fit for this import workflow Boundary that changes the decision Cost-attribution consequence
Sentry Teams that need the fuller error-tracking workflow and frontend ergonomics A scheduled job that never runs still needs a heartbeat signal Attribute usage with project and tenant metadata, then validate the billing model against the team's allocation rules
Bugsnag A dedicated error-tracking product to evaluate alongside Sentry Do not assume product error events prove schedule execution Check whether its project structure maps cleanly to app, team, and tenant ownership
Rollbar Another dedicated error-tracking choice for grouped exceptions and triage evaluation Heartbeat coverage remains a separate acceptance criterion Test the proposed project and environment split before it becomes the chargeback model
Datadog A broader observability suite to evaluate when the importer needs cross-signal investigation Broader scope can exceed an MVP's exception-triage requirement Validate service and tenant tagging against the intended allocation model
Grafana A broader observability approach to evaluate when the team wants to compose several signals The team still owns the design of the scheduled-run signal and page Include the operating cost of the chosen deployment and integrations
Better Stack A managed option to evaluate for a combined operational workflow Confirm that its workflow matches both exception and missing-run acceptance tests Map sources and owning services before rollout
Healthchecks Deadline monitoring for “the job should have run” It complements rather than replaces exception event detail Attribute checks to the owning schedule or service
Simple error API One small app or API that mainly needs capture, grouping, event inspection, search, and resolve The team must build polling and notification routing; there is no trace-query UI, source-map processing, crash symbolication, or replay A narrow API surface is easier to tag, but poller operations and paging delivery are still platform costs

This table is deliberately asymmetric. Healthchecks solves the missing-run half, while the other choices address exception triage to different degrees. Treating them as five interchangeable products would hide the architectural decision.

Infrai fits the simple-API row when consolidating backend services under one key and one bill materially reduces key sprawl and month-end reconciliation. Its error surface covers capture, grouped issues, event payload inspection, search, and resolve, and the broader API provides per-call cost, vendor, and latency metadata.

Infrai's second relevant advantage is a single, unified REST API: 295 routes across 20 modules use consistent conventions, any language or runtime can call them over pure HTTP, and no SDK is required. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages. For this workflow, those traits let the platform team inspect the current error contract and implement the poller without adopting another service-specific client dependency. The trade is explicit: notification routing requires a custom poller over error list or search results, and silent scheduled failures still require a heartbeat tool. That's attractive for a small MVP only when the team accepts ownership of that thin control plane.

The following Go program is the smallest useful core of that poller: it calls the provider's error list, authenticates from the environment, uses an explicit method, honors Retry-After on a 429, and surfaces a non-success body. Set INFRAI_API_BASE_URL to the documented API origin and INFRAI_API_KEY to a scoped key before running it. Production code should persist its last evaluation window and deduplicate the notification after this read.

package main

import (
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    base := strings.TrimRight(os.Getenv("INFRAI_API_BASE_URL"), "/")
    key := os.Getenv("INFRAI_API_KEY")
    if base == "" || key == "" {
        panic("INFRAI_API_BASE_URL and INFRAI_API_KEY are required")
    }

    client := &http.Client{Timeout: 10 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, base+"/v1/errors/list", nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }

        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("error list returned %s: %s", resp.Status, body))
        }

        var result any
        if err := json.Unmarshal(body, &result); err != nil {
            panic(err)
        }
        pretty, err := json.MarshalIndent(result, "", "  ")
        if err != nil {
            panic(err)
        }
        fmt.Println(string(pretty))
        return
    }
    panic("error list remained rate limited after four attempts")
}
Enter fullscreen mode Exit fullscreen mode

The alert-to-action trace

When the page fires, its payload should identify the service, tenant, schedule, expected deadline, last successful run, and a run or error identifier. The first link should land on the evidence that distinguishes “never started” from “started and failed.” A generic “imports are broken” page spends the response budget on navigation.

For a captured exception, the responder searches or opens the grouped issue, inspects the specific event payload, and checks whether other tenants share the signature. Grouping is valuable here, but it is not an incident policy: a single group can contain events with different customer impact, and a newly deployed fix does not prove recovery. Resolve the group after the fix ships, then require a successful scheduled outcome to close the operational incident.

For a missing heartbeat, the responder checks scheduler dispatch and queue handoff before looking for exceptions. There may be none. If trace and span IDs exist in logs, they can help correlate known records, but without distributed trace querying the responder must not promise a service map or span-tree investigation. That limitation becomes expensive as the importer splits across more services, which is a practical migration trigger rather than a theoretical complaint.

The custom poller also needs production semantics. It must persist the last evaluated window, deduplicate notifications, retry delivery without multiplying pages, and report its own health. Route severity by remaining SLO budget: a late run with ample budget may create a ticket, while a missed deadline near breach can page. Phone, SMS, Slack, and webhook routing do not appear by wishing an error list into an alert manager; somebody must own those integrations and their failure modes.

A skeptical buy-versus-build gate

My first design instinct would be to keep the MVP narrow. The correction comes after counting ownership: I would approve the simple path only after the platform team can answer these four questions in writing:

  1. Can one on-call engineer operate the poller, notification integration, and heartbeat checks without displacing higher-value roadmap work?
  2. Can every event and check be attributed to a service, tenant, and schedule without creating a project per customer?
  3. Is exception grouping plus raw event detail enough, or will the product soon require browser source maps, replay, native crash symbolication, or cross-service trace exploration?
  4. What measurable condition triggers migration: service count, on-call toil, detection latency, or an SLO miss caused by missing diagnostic context?
Decision Buy a fuller tracker Build around a simple API
On-call load Prefer when built-in workflows remove recurring platform work Prefer only when the poller and routing remain genuinely small
Lock-in Accept deeper product workflows deliberately Keep the event schema and correlation fields portable
Cost attribution Validate projects, environments, and usage exports against ownership Tag calls and events consistently; include poller and delivery costs
Capability runway Prefer when frontend or trace-oriented investigation is near-term Prefer for one small backend app with limited triage needs

No price figure settles this table. A service that appears inexpensive but adds an unowned alerting component consumes on-call capacity, and a broad suite bought years before its features are needed creates its own operational and procurement weight. Model both. Revisit the decision at a defined threshold rather than after responders have accumulated private scripts.

Thresholds spend attention

The final risk is false precision. Alerting on result_count == 0 catches a useful symptom only when zero is anomalous for that feed; alerting on one exception may page on a retry that completed successfully; waiting for several failures may exhaust the completion SLO. Segment thresholds by schedule and business expectation, keep a record of successful terminal outcomes, and test notification deduplication during deploys.

Every false page spends trust. Every threshold that is too loose spends error budget.

For a small Node.js SaaS MVP, start with the smallest system that covers both axes: grouped exception triage and explicit missing-run detection. Choose a simple error API when the application is small, backend-focused, and the team can own polling and routing. Choose a fuller tracker when frontend diagnostics or mature alert workflows already justify their operational weight. In either case, keep the heartbeat independent enough to tell you that nothing happened.

Further reading

Top comments (0)