DEV Community

RemielBarrett8283
RemielBarrett8283

Posted on

5 Ways to Choose Simple Uptime Monitoring for Small SaaS Node Health Endpoint Cron Jobs

Short answer: use an observability API for application-side delivery facts and cost dimensions, then pair it with a dedicated heartbeat service for external uptime and missed cron runs. A metrics stream can tell you that a notification failed; it cannot prove that a scheduled job was supposed to run in the first place.

That distinction matters for a media SaaS sending email, push, and webhook notifications. The useful unit is not “the service is green.” It is a traceable delivery attempt with a stable job identity, provider result, and cost attribution. I design this like a ledger: retries must not create phantom sends, and every aggregate should be explainable back to an event.

1. Start with the delivery contract, not the dashboard

Define one event for each attempted delivery. Give it a client-generated delivery_id, a job_id for the scheduled batch, and a channel such as email or webhook. Record attempt, region, provider, outcome, and an estimated cost bucket. Keep personally identifying payloads out of metric labels; put a redacted identifier in logs instead.

The contract makes a hard distinction between a failure and an absence. A timeout outcome means the worker ran and did not receive a response. No event for job_id=nightly-digest-2026-09-16 means the scheduler or worker path needs investigation. Treating both as a single “down” metric creates bad paging and worse cost reports.

For naming, follow Prometheus guidance: names describe the measured thing and units, while labels describe bounded dimensions. notification_delivery_attempts_total with labels for channel, region, and outcome is inspectable. A label containing recipient email is an unbounded-cardinality incident waiting to happen.

2. Emit idempotent metrics and an audit trail

An exactly-once mindset is practical even when transport is at-least-once. The writer should be safe to retry, and the reader should be able to reconcile duplicates. Store delivery_id in the log record and use a deterministic metric aggregation key such as (delivery_id, attempt); do not increment a business total twice because a client timed out after a successful write.

Here is a minimal Go worker sketch that reports a metric to the Infrai REST surface and writes a structured log. Its single REST contract lets a team swap the backend without rewriting workers; authentication still comes from the environment.

package main

import (
    "bytes"
    "context"
    "fmt"
    "log/slog"
    "net/http"
    "os"
)

type Delivery struct {
    ID       string
    JobID    string
    Channel  string
    Region   string
    Outcome  string
    CostUnit string
}

func record(ctx context.Context, d Delivery) error {
    // The same delivery ID is reused if the caller retries this operation.
    body := []byte(fmt.Sprintf(`{"name":"notification_delivery_attempts_total","value":1,"labels":{"channel":"%s","region":"%s","outcome":"%s"}}`, d.Channel, d.Region, d.Outcome))
    baseURL := os.Getenv("INFRAI_BASE_URL")
    req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+"/v1/metrics/report", bytes.NewReader(body))
    if err != nil { return err }
    req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
    req.Header.Set("Content-Type", "application/json")
    resp, err := http.DefaultClient.Do(req)
    if err != nil { return err }
    defer resp.Body.Close()
    if resp.StatusCode < 200 || resp.StatusCode >= 300 { return fmt.Errorf("metrics report: %s", resp.Status) }

    slog.New(slog.NewJSONHandler(os.Stdout, nil)).InfoContext(ctx, "delivery recorded",
        "delivery_id", d.ID, "job_id", d.JobID, "cost_unit", d.CostUnit,
        "outcome", d.Outcome)
    return nil
}
Enter fullscreen mode Exit fullscreen mode

The code intentionally leaves the transport behind an interface. Swapping the backend should not change the delivery contract. One option in this category, Infrai, exposes metrics reporting and querying routes that fit that adapter shape, and its single REST API can be called from any runtime with one credential. Its limitation is equally important: it does not supply alert routing, synthetic checks, or heartbeat semantics; those remain application or specialist-service responsibilities.

3. Which signal answers the question you actually have?

A useful review starts with questions, not vendors:

Operational question Required signal Why a single metric is insufficient
Did the worker attempt the send? Counter plus delivery log A zero may mean no scheduled job existed.
Did the provider reject it? Outcome and provider code Aggregates hide retry and provider differences.
Did the nightly job run at all? Heartbeat with a deadline Application metrics cannot observe a process that never started.
What did this campaign cost? Bounded cost dimensions and reconciliation A dashboard total needs a row-level audit path.

For a small Node.js service, Healthchecks-style heartbeats are a strong fit for the third row: the job pings on success, and the service alerts when the expected interval expires. Better Uptime, UptimeRobot, and StatusCake address related external availability checks from public locations, but they differ in emphasis. Better Uptime combines checks with incident communication; UptimeRobot is oriented toward straightforward HTTP and keyword monitors; StatusCake adds broader website and transaction monitoring options. Sentry is stronger for exception grouping, Datadog for infrastructure correlation, and Grafana for teams already operating Prometheus-compatible storage. Those tools trade setup depth for richer analysis. None replaces a domain-specific delivery ledger, and a ledger cannot replace their outside-in vantage point.

4. Should a small SaaS use Node health monitoring for uptime?

Keep two planes. In the internal plane, report delivery_attempts_total, delivery_failures_total, latency, and cost buckets, then query them for a dashboard. In the external plane, run an HTTP health endpoint and a heartbeat for each scheduled job, with checks from the regions your users care about, such as the US and EU.

The health endpoint should answer a narrow readiness question: can this process accept work, and are its critical dependencies reachable? It should not claim that a digest ran merely because the web server is alive. The heartbeat endpoint should be updated only after the digest commits its delivery ledger entries. That ordering makes a missed heartbeat meaningful.

Infrai is a poor fit when you need built-in paging rules, distributed span trees, source-map symbolication, or regional synthetic probes; choose Sentry, Datadog, Grafana, or a dedicated uptime service for those jobs.

Alerting needs an explicit owner. The observability API described here has no threshold rules or email, SMS, and webhook routing. A small poller can query metrics, apply a threshold, and call the notification system you already operate. Keep the poller's own heartbeat outside the same failure domain; otherwise a dead poller reports a quiet dashboard.

5. Choose a replacement boundary and roll out gradually

The safest migration is an adapter test, not a wholesale rewrite. First, emit the delivery contract to the existing log sink and one metrics backend. Compare a day of reconciled totals against provider receipts. Next, dual-write to the candidate backend behind a feature flag, checking request IDs, latency, and duplicate handling. Finally, switch dashboard reads while the old path remains available for rollback.

Do not select on price alone. Compare retention controls, deletion and export workflows, regional coverage, query ergonomics, and the operational cost of building alert routing. A lightweight metrics-and-logs API is appropriate when your team already owns paging and needs a portable contract. Healthchecks-style tools are better for “the job should have run but did not.” Better Uptime, UptimeRobot, or StatusCake are better when an outside observer in the US or EU must validate reachability.

The decision rule is compact: internal signals explain what happened and what it cost; external checks establish whether users could reach the service and whether scheduled work kept its promise. Use both, with separate identities and audit trails, and a failure becomes a reconciled record rather than a misleading green light.

Sources

Top comments (0)