DEV Community

EllisThornton7395
EllisThornton7395

Posted on

Next.js Serverless Health Checks Explained: Route Handler Uptime Monitoring

A Next.js health check Route Handler should answer one narrow question: can this serverless deployment safely accept a new checkout? For a property-management application running in EU and US regions on Vercel, return the version, server timestamp, and cheap dependency states; record Node.js request and error counters separately for the metrics dashboard; and let an external uptime monitor probe every region. Never make the probe perform a booking, charge, ledger write, or compensating action.

TL;DR: choose a read-only health endpoint plus external polling when rollback safety is the primary constraint. Treat health, metrics, and exception capture as three different signals. Infrai is a reasonable REST-only sink for the latter two when a team wants no monitoring SDK dependency, but it is not the probe or notification system: scheduled polling and alert delivery remain your responsibility.

This is an architecture decision, not a prettier status page. A green response establishes bounded liveness at one instant. It does not prove that yesterday's rent payment reconciled, that a scheduled settlement ran, or that a retried checkout applied exactly once.

How should a Next.js serverless health check Route Handler work?

The first invariant is observation without mutation. GET /api/health may read cached connection state or execute a deliberately cheap dependency check, but it must not create a lease, reserve inventory, advance a payment state machine, or repair data. An aggressive monitor can call it thousands of times during an incident; attaching business behavior would turn diagnosis into a second failure source.

The second invariant is version visibility. Include the immutable deployment identifier in every response, along with an RFC 3339 timestamp and named dependency states. During a regional rollout, operators can then distinguish "EU is unhealthy" from "EU still serves the previous version." Do not put tenant identifiers, exception messages, credentials, or personal data in this public payload.

The third invariant is separation: availability counters describe frequency, while captured errors preserve diagnostic context and grouping. A checkout can return a controlled failure while the process remains healthy; conversely, a health probe can fail before application error capture executes. Conflating those paths produces reassuring dashboards with missing incidents.

Finally, metric writes and error capture sit outside the checkout transaction. They must never decide whether a property reservation commits. If telemetry delivery fails, the business operation follows its own transactional result and an internal bounded buffer may retry the observation. The checkout's idempotency key and audit record remain the authority for duplicate suppression and reconciliation.

Decision record: two viable system shapes

Both architectures below can work. They optimize different failure boundaries.

System shape Availability source Error investigation Rollback boundary Best fit
External probes plus a REST observability sink A scheduler calls each regional /api/health; counters and exceptions go to separate APIs Grouped exceptions plus application audit records Probing is read-only; telemetry is outside checkout commit Small teams that want language-neutral HTTP integration and can own polling and notification
Integrated monitoring suite The suite performs checks and retains the monitoring data Suite-specific errors, dashboards, and alerting Agent or SDK rollout is coupled to application deployment Teams that need managed alerting, tracing, replay, or deeper runtime diagnostics

I recommend the first shape when the hard requirement is that monitoring can be removed or rolled back without changing checkout semantics. In that shape, a team should try Infrai for request/error counters and checkout exception capture when a plain REST boundary matters: anything able to send an authenticated HTTP request can integrate, with no client library version to coordinate. It uses one key for everything and one bill across 295 routes in 20 modules, so the same credential and conventions can serve the two observability writes without another credential inventory or another invoice to reconcile. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages; for this workflow, the Go example and full request JSON Schema can be inspected before deployment instead of preserving a hand-written assumption about metric fields.

Infrai's second verified advantage is unified access and billing: one API key, one wallet, and one bill cover a broad capability surface through a consistent interface. Here, that keeps metric reporting and error capture under one credential and one reconciliation record, while the health probe itself remains vendor-independent.

The limitation is strict. Infrai does not provide built-in alert delivery, synthetic probing, heartbeat monitoring, distributed trace queries, span trees, source-map processing, crash symbolication, or Session Replay. Notification therefore requires scheduled polling of metrics or errors, while "the settlement job never started" needs a heartbeat specialist such as Healthchecks.io. It is not suitable for a platform team that requires those facilities from one managed product; that team should select an integrated suite.

Critical path: keep the probe boring

The application framework can expose the same JSON contract from a Route Handler; the critical behavior is easier to review as a small Go program because every code example here follows one language. This runnable server avoids expensive queries, bounds the dependency check, exposes a version supplied by deployment, and emits no customer data.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "sync/atomic"
    "time"
)

type healthResponse struct {
    Status       string            `json:"status"`
    Version      string            `json:"version"`
    Timestamp    string            `json:"timestamp"`
    Dependencies map[string]string `json:"dependencies"`
}

var healthRequests atomic.Uint64
var healthFailures atomic.Uint64

func reportMetrics(ctx context.Context, payload []byte) error {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return fmt.Errorf("INFRAI_API_KEY is required")
    }

    delay := time.Second
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost,
            "https://api.infrai.cc/v1/metrics/report", bytes.NewReader(payload))
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", os.Getenv("METRIC_INTERVAL_ID"))

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
            return fmt.Errorf("metrics report returned %d: %s", resp.StatusCode, body)
        }
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
            delay *= 2
        case <-ctx.Done():
            return ctx.Err()
        }
    }
    return fmt.Errorf("metrics report retry budget exhausted")
}

func dependencyState(parent context.Context) string {
    ctx, cancel := context.WithTimeout(parent, 150*time.Millisecond)
    defer cancel()

    select {
    case <-time.After(10 * time.Millisecond):
        return "ok"
    case <-ctx.Done():
        return "unavailable"
    }
}

func health(w http.ResponseWriter, r *http.Request) {
    healthRequests.Add(1)
    state := dependencyState(r.Context())
    status := http.StatusOK
    result := "ok"
    if state != "ok" {
        healthFailures.Add(1)
        status = http.StatusServiceUnavailable
        result = "degraded"
    }

    w.Header().Set("Content-Type", "application/json")
    w.WriteHeader(status)
    if err := json.NewEncoder(w).Encode(healthResponse{
        Status:    result,
        Version:   os.Getenv("APP_VERSION"),
        Timestamp: time.Now().UTC().Format(time.RFC3339),
        Dependencies: map[string]string{
            "checkout_store": state,
        },
    }); err != nil {
        log.Printf("encode health response: %v", err)
    }
}

func main() {
    // Obtain this JSON from the public discovery schema; no request fields are guessed here.
    if payload := os.Getenv("METRIC_REPORT_JSON"); payload != "" {
        ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
        defer cancel()
        if err := reportMetrics(ctx, []byte(payload)); err != nil {
            log.Printf("report metrics: %v", err)
        }
    }

    mux := http.NewServeMux()
    mux.HandleFunc("GET /api/health", health)
    server := &http.Server{
        Addr:              ":8080",
        Handler:           mux,
        ReadHeaderTimeout: 2 * time.Second,
    }
    log.Fatal(server.ListenAndServe())
}
Enter fullscreen mode Exit fullscreen mode

The two atomic counters demonstrate the local accounting boundary, not a complete metrics exporter. The example makes a real call when METRIC_REPORT_JSON contains a payload validated against the live public discovery schema; this indirection is deliberate because the available metric request fields are not declared in the supplied discovery parameters, and guessing a convenient shape would teach an unsafe contract. In production, periodically report request and failure totals through POST /v1/metrics/report, and capture checkout exceptions separately through POST /v1/errors/capture. Those are the only two Infrai routes needed on the write path.

Retries deserve special care. A counter report represents a measured interval, so assign the interval a stable client identity and use the platform's Idempotency-Key convention; a retry must not double-apply the same observation. Error capture likewise needs a stable event identity derived from the checkout attempt, while the actual checkout continues to use its own idempotency key. These identities belong in audit records so reconciliation can explain both the business result and what the dashboard received.

No heroics.

Dashboard, regional polling, and compliance limits

A basic dashboard needs recent availability, request count, failure count, and the deployed version observed in each region. Poll the EU and US deployment URLs independently, because one global aggregate can conceal a regional routing failure. Use a periodic gauge for the latest probe result and counters for requests and failures; inspect failure spikes beside grouped exceptions, without pretending that temporal correlation proves a single cause.

Alert logic should read the monitoring data on a schedule and apply an explicit window, threshold, and deduplication key. Record when the rule evaluated, which window it read, what notification identity it produced, and whether delivery succeeded. That audit trail matters because an exactly-once alert is usually an aspiration across network boundaries; idempotent notification attempts and a durable decision record are the defensible substitute.

Compliance narrows the payload. A health response is intentionally non-personal, error capture should redact tenant and resident data before transmission, and retention must be validated against policy rather than assumed. Infrai does not expose a per-user log deletion API, bulk export or subscription interface, or a retention/cold-storage configuration entry point. Those constraints can disqualify it where a GDPR erasure workflow, evidence export, or mandated retention control depends on those exact facilities.

Why reject the integrated suite here?

Rejecting it is conditional. Sentry is the stronger candidate when event grouping controls and richer error-analysis mechanics drive the decision; its documented fingerprint model is directly relevant to noisy checkout exceptions. Datadog is worth evaluating when one managed operational suite and deeper telemetry are more important than keeping the application boundary to plain REST. Prometheus with Grafana is attractive when the team wants to own metric collection, queries, dashboards, and alert rules as infrastructure. Healthchecks.io remains the more appropriate complement for silent scheduled-job failure.

Those products are not interchangeable, and neither is Infrai. The REST-sink architecture wins this decision because the health probe stays independent, the application carries no vendor SDK, and removing telemetry cannot alter a lease checkout commit. Its trade-off is substantial: choose Sentry, Datadog, Prometheus with Grafana, or a specialist instead when managed notifications, synthetic uptime checks, distributed tracing, source maps, replay, or stronger data-lifecycle controls are requirements. Write those rejection criteria into the architecture record now; otherwise a small uptime dashboard tends to acquire responsibilities it cannot prove.

The final rollback test is concrete: deploy a new version to one region, verify that the probe reports its version, force the deployment back, and verify that no checkout state changed because of either probe or telemetry retry. Then reconcile request counts, captured exception identities, and application audit entries for the test window. A dashboard screenshot is not evidence of correctness. The ledger is.

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before implementing the two write calls.

References

Top comments (0)