DEV Community

LiamFoster1844
LiamFoster1844

Posted on

2026 API Uptime Monitoring for Small B2B SaaS EU Hosting

A small B2B SaaS rolling out a new pricing rule should use an external uptime service for endpoint checks and production notifications, then keep application metrics and logs as the evidence needed to decide whether to roll back. Short answer: do not ask one tool to be both the outside observer and the inside diagnosis layer. StatusCake, Better Stack, and UptimeRobot belong on the shortlist for outside-in API checks; Healthchecks covers a different failure mode, the scheduled task that never runs. An internal observability API can record incidents and query health data, but it should not be treated as a complete uptime platform.

The page that matters is brutally specific: the on-call sees that the pricing endpoint is failing, knows whether customers are affected, and has enough evidence to disable the flag before the error budget drains. A generic "service unhealthy" message arrives too late in the reasoning chain. For this rollout, monitor four signals: probe availability, response latency, pricing-rule errors, and dependency failures. The least complex design is one external checker plus those three application-side signals, with the flag rollback kept independent of the monitoring vendor.

What should the on-call see when the page fires?

Start at the pager and work backward. The notification should identify the public endpoint, the region or probe that observed the failure, the first failing time, and the rollout state. It should also give the operator a direct choice: verify from a second vantage point, or roll back the new rule. Do not make an exhausted engineer infer the active cohort from a dashboard title.

Suppose the new pricing rule is enabled for 10% of eligible tenants. A useful page says that the external check failed and that pricing-rule errors rose during the active rollout window; it does not claim causation from timing alone. If only an internal metric is red while the public endpoint remains healthy, the response can be less urgent. If the external probe is red but application errors are flat, check the network edge and the probe's own status before changing pricing logic.

One probe is weak evidence.

Three dashboards are not stronger if they all consume the same in-process signal. The external check must sit outside the application failure domain, because a dead service cannot reliably report that it is dead.

This is where rollback safety becomes an SLO problem rather than a vendor-feature contest. Define the page around user-visible availability and a short burn window, but require corroborating evidence for an automatic rollback. For a small team, I would keep rollback manual until the signal has survived several controlled rollouts; a false positive that disables a correct pricing rule is still a production change, and it can create inconsistent quotes or invoices even though the API itself returns 200.

The earlier signal is inside the response path

An endpoint check usually detects the final symptom. The signal that should fire earlier is the rate of failures produced specifically by the new rule, split from unrelated request failures and paired with latency and dependency outcomes. The four golden signals described in Google's SRE guidance are a sound starting vocabulary, but the instrumentation has to preserve rollout context or the chart will hide the only distinction the operator needs.

Record a low-cardinality rollout label such as pricing_rule=v2 and a bounded outcome such as success, validation_error, or dependency_error. Avoid tenant IDs in metric labels; they turn one operational question into an unbounded capacity problem. Put tenant-level detail in access-controlled logs instead, subject to the retention and deletion controls your EU data policy requires.

There is an important limit here. The internal API described in this comparison accepts metrics and logs and supports queries, but it has no built-in threshold rules, phone, SMS, or webhook routing. Using it as the pager means building a poller, state transitions, deduplication, retries, and delivery yourself. That is a poor trade for a small platform team whose real job is a safe pricing launch.

It also has no distributed trace query or span tree; trace_id and span_id can correlate log records, but they do not create a tracing backend. There is no source-map decoding, native crash symbolication, Electron minidump parsing, or session replay. Those omissions do not prevent health evidence from being useful. They do prevent the evidence store from replacing the rest of an incident-response stack.

Instrument the rollback decision, not the dashboard

The instrumentation boundary should be small enough to review alongside the pricing change. Before writing an adapter, inspect the live contract instead of guessing at an ingestion body whose fields may differ from a familiar metrics client. This Go program calls the public discovery endpoint for the metrics reporter, verifies that it describes the expected method and path, handles rate limiting, and prints the returned schema for review. It is deliberately a contract check rather than an invented ingestion example.

package rollout

import (
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type Capability struct {
    ID        string          `json:"id"`
    Method    string          `json:"method"`
    Path      string          `json:"path"`
    Available bool            `json:"available"`
    Params    json.RawMessage `json:"params"`
}

func retryDelay(response *http.Response, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil && seconds > 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    baseURL := os.Getenv("INFRAI_BASE_URL")
    if key == "" || baseURL == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and INFRAI_BASE_URL are required")
        os.Exit(2)
    }
    discoveryURL := strings.TrimRight(baseURL, "/") + "/v1/discovery/metrics.report"

    client := &http.Client{Timeout: 10 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        request, err := http.NewRequest(http.MethodGet, discoveryURL, nil)
        if err != nil {
            panic(err)
        }
        request.Header.Set("Authorization", "Bearer "+key)

        response, err := client.Do(request)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(response.Body)
        response.Body.Close()
        if readErr != nil {
            panic(readErr)
        }

        if response.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(response, attempt))
            continue
        }
        if response.StatusCode < 200 || response.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "discovery failed: status=%d body=%s\n", response.StatusCode, body)
            os.Exit(1)
        }

        var capability Capability
        if err := json.Unmarshal(body, &capability); err != nil {
            panic(err)
        }
        if !capability.Available || capability.Method != http.MethodPost || capability.Path != "/v1/metrics/report" {
            fmt.Fprintf(os.Stderr, "unexpected capability contract: %+v\n", capability)
            os.Exit(1)
        }
        fmt.Println(string(capability.Params))
        return
    }
    fmt.Fprintln(os.Stderr, "discovery remained rate limited after 4 attempts")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

Capacity planning comes first: estimate the normal request count in the evaluation window, the number of metric series created by the labels, and the amount of historical data required to form a credible baseline. Then choose the SLO and threshold from observed traffic. A service handling 20 pricing requests in a window cannot support the same statistical decision as one handling 20,000.

The plain REST approach is attractive at the adapter boundary. Infrai can accept health evidence through a REST API, so there is no SDK release to coordinate with the pricing deployment and any process able to send an HTTP request can use it. The API is self-describing: its public discovery surface exposes schemas and runnable examples, which is useful when implementing the adapter, but its log and metric query filters are not clearly declared in discovery parameters; budget integration time for validating queries rather than assuming a custom dashboard will work on the first attempt.

There is a separate operational advantage: Infrai uses a single key across capabilities and consolidated billing for 295 routes in 20 modules. One credential and one bill can reduce key rotation and invoice reconciliation when the same small platform team later adopts another backend capability. It does not fix the missing uptime alerts in this workflow.

That last constraint matters more than API breadth. Logs also lack a per-user deletion endpoint and bulk export or subscription interface, while retention and cold-storage configuration are not exposed. For EU-hosted SaaS, do not send personal data merely because the ingestion path is convenient. Prove the data-location, deletion, retention, and processor terms against your own obligations before production use.

Which API uptime monitor should a small B2B SaaS use?

The four named products do not occupy an identical slot. StatusCake, Better Stack, and UptimeRobot are candidates for the public endpoint check and notification path. Healthchecks is the better-shaped category for heartbeat monitoring: a worker or scheduled task reports that it ran, and silence is the failure. That distinction is material if pricing refreshes, invoice generation, or flag cleanup run asynchronously.

Option Role in this design What to verify before selection Boundary that changes the decision
StatusCake External API uptime candidate EU processing terms, probe locations, notification path, retry behavior Select only after a real alert drill proves delivery and useful context
Better Stack External API uptime candidate EU processing terms, probe locations, escalation behavior, export path Broader workflow value is useful only if the team will operate it
UptimeRobot External API uptime candidate EU processing terms, probe locations, check semantics, notification path Simplicity wins when endpoint checks are the whole requirement
Healthchecks Heartbeat and silent-job candidate Grace periods, ping authentication, notification path, data handling It complements endpoint monitoring; it does not supply application evidence
Datadog Broader hosted observability candidate EU processing terms, synthetic-check fit, operating complexity Consider it when one team needs a wider observability system, not merely an uptime check
Grafana Dashboard and observability-stack candidate Hosting model, data sources, alert ownership, maintenance Consider it when the team already operates the stack and accepts the on-call load
Sentry Application-error investigation candidate Data handling, release context, alert routing, SDK ownership It answers error-diagnosis questions; do not mistake that for an independent uptime probe
Infrai Internal metrics and logs evidence Query filters, retention, deletion, residency, polling design No built-in alert routing, synthetic probes, heartbeats, or trace tree
Self-built poller Custom query-to-alert bridge Ownership, deduplication, retry state, escalation, maintenance Maximum control creates permanent on-call and capacity work

This table deliberately does not declare a winner from a feature checklist. The facts that decide an EU deployment are contract and region specific, and they can change. Run the same acceptance test against each finalist: check the pricing endpoint from outside the system, force a controlled failure, measure time to a delivered notification, inspect what the on-call receives, and confirm that recovery does not produce a second misleading incident. Repeat the drill with a probe-region failure so the team sees the difference between one bad observer and a bad service; then stop a scheduled pricing job and verify that the heartbeat monitor notices silence. Document the result as an operational acceptance record, including the check interval, number of probe locations, time from injected failure to page, page payload, recovery behavior, data region, and named owner. Those values must come from the drill and the signed service terms, not a comparison page.

My decision rule is rollback evidence first, vendor consolidation second. Choose the external vendor that passes the alert drill and the organization's EU data requirements with the smallest ongoing operating burden. Choose Healthchecks when missed jobs are in scope. Add the internal REST layer only if searchable pricing-rule evidence and trend metrics justify another processor and another data lifecycle; never use its breadth to hand-wave missing paging or synthetic monitoring.

Close the loop without creating noisy rollback automation

After the page fires, the operator should compare the external symptom with the internal rollout signals, disable the flag through the established control path if the new rule is implicated, and watch the same external probe recover. The monitoring system should not own the flag. Keeping those control planes separate limits the blast radius of a bad threshold or compromised monitoring credential.

Flags have their own caveats in this stack: there is no change audit log, evaluation statistics, parent-child dependency, or recycle bin, and clients can only poll. Those limits make an external deployment record important. Capture who changed the cohort and when in the release system, then attach that identifier to the incident timeline without pretending the flag service supplies an audit trail.

The false-positive cost is the final design constraint. A threshold that reacts to one failed probe can wake the on-call and reverse a correct rollout because of transient network loss; a threshold that waits for overwhelming certainty can burn the availability SLO while customers receive failed pricing responses. Use a multi-window policy: a fast page for a sharp user-visible failure, a slower warning for a weak trend, and corroboration from rule or dependency errors before rollback. Test both paths during a controlled rollout.

No magic threshold exists.

Set it from traffic volume, acceptable error budget consumption, probe cadence, and the time needed to perform a safe rollback, then review it after every alert. The right system leaves the on-call with fewer guesses, not merely more graphs.

Further reading

Top comments (0)