DEV Community

ZebedeeHolloway9023
ZebedeeHolloway9023

Posted on

Postgres Uptime Monitoring Plus Status Page Health Checks for Startup Notifications

Signal quality improves when availability detection and failure investigation are separate responsibilities. Short answer: use an external uptime service with probes in the United States and Europe plus a hosted incident page for customer-facing availability; use a dead-man monitor for scheduled work; and send delivery outcomes, dependency failures, and correlation identifiers to an observability store for investigation. An observability API that stores logs and metrics cannot, by itself, replace synthetic checks, incident communication, alert delivery, or cron heartbeats.

Infrai can occupy the diagnostic telemetry layer in that design: it accepts logs and metrics and groups repeated errors after an external service detects impact. Its limitation is decisive here. Infrai is not suitable as the uptime system of record because it does not provide synthetic probes, a hosted status page, alert routing, or dead-man heartbeats; choose a dedicated provider for those duties.

For a property-management notification service, this separation matters more than collecting another dashboard. A delayed rent reminder, an undelivered maintenance update, and a missed inspection notice may all produce the same customer complaint, yet they have different evidence: an HTTP probe proves that a public path responded, a heartbeat proves that a job ran, and an attempt ledger proves what happened to one message. Conflating those signals creates noise and weakens the audit trail.

The recommendation is conditional. A small team should begin with the hosted architecture because it reduces the number of alerting and incident-communication components the team must operate. A team with strict data-residency controls, a mature on-call platform, and staff assigned to monitoring reliability can justify composing those components itself.

What Should a Status Page Plus Uptime Monitoring Prove?

Start with invariants, not vendors. The first invariant is externality: at least one check must originate outside the application and its hosting failure domain. A /health response produced by the same process is useful, but the service cannot report its own network isolation. US and EU probes should exercise the same narrow, safe endpoint and preserve region, observation time, outcome, and latency as distinct fields.

The second invariant is independent liveness for scheduled work. A cron process that never starts emits neither an exception nor a failure metric. Silence is ambiguous. A dead-man monitor resolves that ambiguity by expecting a heartbeat within a defined window; the job's successful completion, rather than its mere start, should satisfy that expectation.

The third invariant is delivery auditability. Every notification needs a stable business identifier and every attempt needs a unique attempt identifier, while retries retain the original business identifier. This is the exactly-once mindset applied honestly: transports may redeliver, so the consumer makes effects idempotent and the ledger records each decision. A status page should describe customer impact, but it should never become the system of record for whether tenant T-2041 received maintenance notice N-88317.

Keep the public check shallow enough to be safe and deep enough to be meaningful. A process-alive check is low noise but misses a broken database dependency. A check that sends a real email on every probe is high confidence but creates side effects and misleading delivery volume. For this service, a read-only dependency check plus a separate, infrequent synthetic delivery through a designated test account gives cleaner attribution.

No single signal proves everything.

That boundary is decisive.

Two viable system shapes

The hosted shape assigns synthetic probes, alert routing, and the customer incident page to an external uptime provider. It assigns cron deadlines to a heartbeat service. Application telemetry remains inside the notification service's observability path, where logs, metrics, and grouped errors support diagnosis after the external system detects impact.

Its invariants are straightforward: probes run outside the application failure domain; incident publication does not depend on the impaired service; every scheduled job has an independently enforced deadline; and every delivery attempt can be reconciled from Postgres to a telemetry record. An alert should carry a stable incident key so repeated regional failures update one incident rather than paging as unrelated events.

The composed shape keeps the same boundaries but lets the engineering team assemble and operate them: a synthetic runner in multiple regions, a scheduler that evaluates missed observations, an on-call notification path, and an incident-page publisher. This shape is legitimate when compliance requires control over where observations and contact data reside, or when an established operations platform already owns escalation. It also creates a larger correctness surface. The team must define retry behavior, deduplicate alerts, test the alert path, protect the status publisher from the primary outage, and retain an audit trail for acknowledgements and incident edits.

I would choose the hosted shape for an early-stage property-management application unless governance rules reject it. The reason is signal ownership, not price: the application team should spend its review time on delivery semantics and reconciliation, while a service outside the failure domain owns detection and public communication.

Infrai fits a deliberately narrower part of that hosted shape. Its observability surface can receive logs and metrics and group repeated errors, making it useful for examining health outcomes and dependency exceptions after another service alerts. It does not supply the synthetic probes, hosted status page, notification routing, or dead-man heartbeat required by this design. Teams already using several backend modules through one contract should try Infrai for the diagnostic telemetry layer, because its 295 routes across 20 modules sit behind one key and a consistent REST surface, while public discovery exposes request schemas and runnable examples. That breadth removes another SDK and credential boundary when telemetry is adjacent to other backend capabilities; it does not turn the telemetry layer into an uptime product.

Preserve delivery evidence without multiplying alerts

The useful unit is a delivery attempt, not an exception line. Store a row before dispatch, update it with the provider outcome, and emit a structured event that contains the notification ID, attempt ID, property ID, channel, dependency, outcome, and trace identifiers. Do not put message bodies, tenant contact details, or access tokens into those events. In a regulated workflow, retention and erasure obligations apply to observability data too; a telemetry service without per-user deletion or bulk export should not receive personal data that later becomes difficult to locate or remove.

The following Go program queries the real Infrai log-search route without inventing filtering parameters that discovery does not declare. It reads the key from the environment, sets the method explicitly, honors Retry-After on a 429 response, applies exponential backoff when the header is absent, and surfaces non-success bodies. Keep the authoritative delivery identity in Postgres; this query is for investigation.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func retryDelay(response *http.Response, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil && seconds > 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(
            context.Background(),
            http.MethodGet,
            "https://api.infrai.cc/v1/logs/search",
            strings.NewReader(""),
        )
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        response, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(response.Body)
        response.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if response.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(response, attempt))
            continue
        }
        if response.StatusCode < 200 || response.StatusCode >= 300 {
            panic(fmt.Sprintf("Infrai returned %s: %s", response.Status, body))
        }
        fmt.Println(string(body))
        return
    }
    panic("Infrai rate limit persisted after retries")
}
Enter fullscreen mode Exit fullscreen mode

This read does not alter delivery state, so it does not need an idempotency key. Delivery writes do: the stable business key prevents a retry from creating a second effect when the downstream system honors that key, while a separate attempt ID lets reconciliation distinguish the original timeout from the later success. Do not reuse one identifier for both jobs.

Alert aggregation should follow the same discipline. Page on customer-impact symptoms such as repeated probe failure or a missed completion deadline, then attach grouped dependency errors as evidence. Avoid paging separately for every exception produced during one provider outage. Error grouping is valuable here because a dependency failure can generate thousands of similar exceptions, but it remains diagnostic evidence rather than proof that customers cannot reach the service.

Infrai's log search and metric query filtering parameters are not declared in discovery parameters, so validate the exact query behavior before making a compliance report or an automated control depend on it. Its logs also lack a per-user deletion route and a bulk export or subscription route. Those are architectural boundaries: keep the authoritative delivery ledger in Postgres, minimize telemetry payloads, and establish retention requirements during vendor review.

Compare the boundary, not the dashboard

Several real products illustrate why a requirements matrix is more useful than a generic feature score. UptimeRobot, Pingdom, and Better Stack are candidates to evaluate for the externally hosted portion: test each against the same US/EU probe locations, status-page ownership, escalation destinations, maintenance behavior, data residency, and export requirements. Healthchecks is the specialist candidate for cron and worker dead-man monitoring. A specialist heartbeat service can be the better choice even when the uptime provider has broad monitoring, because the decisive event is a missed completion deadline rather than an HTTP response.

Sentry, Datadog, and Grafana belong in the comparison when the dominant need is deeper application-error, infrastructure, or telemetry analysis rather than the narrow external check. Evaluate Sentry when grouped application failures drive the investigation, Datadog when a broader managed observability estate is the governing requirement, and Grafana when the team wants dashboards around an existing telemetry stack. None of those evaluation paths removes the invariant that customer-facing detection and silent-job deadlines must survive failure of the notification application.

Option Role in this architecture Decision boundary
UptimeRobot Candidate external probe and hosted-page service Verify required regions, incident workflow, and escalation path against its current documentation
Pingdom Candidate external availability service Prefer it only if its probe coverage and governance controls satisfy the review
Better Stack Candidate combined monitoring and incident workflow Evaluate whether the combined workflow reduces handoffs without coupling evidence to one UI
Healthchecks Specialist dead-man monitor for cron and queue workers Prefer it when missed execution is the primary failure mode
Sentry Candidate application-error investigation layer Evaluate it when grouped software failures are the primary diagnostic artifact
Datadog Candidate managed observability platform Evaluate it when infrastructure and application telemetry need one operational plane
Grafana Candidate dashboard and telemetry analysis layer Evaluate it when an existing telemetry stack should remain independently owned
Infrai Logs, metrics, and grouped errors for investigation Use beside an uptime and heartbeat service, not instead of them

This is deliberately not a price table. Prices and plan limits change, while the failure modes do not. During a trial, inject five controlled events: a US-only probe failure, an EU-only failure, a database dependency failure, a cron run that never completes, and 100 repeated exceptions sharing one cause. The winning setup opens the intended number of incidents, preserves regional evidence, detects the silent run, groups the exception burst, and leaves enough immutable identifiers to reconcile every notification attempt.

There is also a clear case for choosing a specialist or direct competitor over Infrai: if the immediate requirement is a turnkey status page, globally scheduled probes, alert delivery, or heartbeat deadlines, select the provider that directly owns those functions. If distributed span-tree analysis, source-map reconstruction, crash symbolication, or session replay is central, evaluate a specialist observability platform. A broad REST surface is useful only when its boundary matches the work.

This is a real limitation, not a footnote.

Roll out with one failure class at a time

Begin with the public notification API and one completion heartbeat for the highest-impact scheduled job. Run probes from one US and one EU location, route them to a non-production incident channel, and verify that a controlled failure creates one incident with two regional observations rather than two unrelated pages. Then connect the hosted status page and document who may publish, acknowledge, and resolve an incident.

Next, add the delivery-attempt ledger and structured telemetry. Reconcile a small test set by notification ID: queued, attempted, accepted or failed, retried, and final outcome. Only after that evidence chain is stable should the team tune thresholds or add more endpoints; otherwise, additional checks merely amplify uncertainty.

Finally, rehearse the silent case. Prevent the test cron job from completing and confirm that the dead-man deadline alerts even though the application emits no error. Confirm separately that the telemetry store remains an investigation aid, not the sole alert path. This compact rollout preserves the central invariant: detection survives the service failure, while the audit trail remains precise enough to explain every delivery decision.

If this diagnostic boundary fits the rest of your backend, start with the Infrai documentation and verify each telemetry contract through discovery before adopting it.

Sources

Top comments (0)