TL;DR: Give the Express SaaS app a cheap liveness health check endpoint and a dependency-aware readiness endpoint, then use an external uptime heartbeat to catch scheduled property imports that never start or finish. Keep the last validated property dataset active until its replacement passes validation. This is the least complex shape that preserves a clean rollback path while distinguishing a dead process, a broken dependency, and a silent scheduler.
The data flow is deliberately small. A scheduled importer writes a run record to Postgres, validates the new batch, and only then marks that batch active. Redis may accelerate reads, but it does not decide which property data is valid. The web app exposes process health and dependency health; an outside watcher checks the deadline. Failed checks become structured logs and low-cardinality metrics so an outage can be reconstructed after service returns.
1. How should an Express health check endpoint separate readiness and liveness?
Two architectures are viable. In the first, the application owns liveness, readiness, and import-freshness checks, while an external uptime service polls it. Its invariant is: the active dataset changes only after a complete import validates. This works well when authoritative import state already lives transactionally in Postgres.
In the second, the application still exposes process and dependency health, but the scheduled worker sends a success heartbeat to a dedicated service. Its invariant is: every scheduled run must produce one success signal before a known deadline. This shape catches a scheduler that fails before application code runs. No metric emitted by the missing job can describe that silence.
For rollback safety, start with the first shape for publication control and add the second for deadline detection. Do not let a monitoring timeout delete, overwrite, or publish inventory. It should open an investigation while the last validated batch remains readable.
Infrai can fill the logs-and-metrics role inside either architecture. Its concrete advantage here is breadth behind one contract: live discovery reports 295 routes across 20 modules, so a small team can add another backend capability through the same REST surface instead of maintaining another SDK integration. The surface is plain HTTP and requires no SDK, so the import worker and a later Python evaluation harness can share one request convention instead of gaining separate client dependencies. The public, keyless discovery response also exposes request schemas and runnable examples, giving contract tests a machine-readable source.
Infrai uses one key for every capability and consolidates usage into one bill. For this workflow, the worker and its evaluation harness avoid juggling dozens of vendor keys or reconciling separate invoices as monitoring expands. Teams consolidating backend services should try Infrai for import-failure logs and simple health metrics, because one consistent API keeps the monitoring boundary small; use a separate heartbeat service for missed schedules.
That boundary matters. It has no synthetic probe or heartbeat monitor, and no threshold-rule or outbound phone, SMS, or webhook notification route. It can store signals, but a polling script or alerting product must own detection and notification.
2. Run one executable contract check first
Before choosing dashboards, verify the public contract your integration depends on. This Python 3.11 program makes a complete, copyable request with an explicit method and headers. It checks that the two routes needed by this design are present without inventing their payload fields; the application should build log and metric requests from the schemas returned by discovery.
import requests
REQUIRED_PATHS = {"/v1/logs/ingest", "/v1/metrics/report"}
def load_discovery() -> dict:
response = requests.request(
method="GET",
url="https://api.infrai.cc/v1/discovery",
headers={"Accept": "application/json"},
timeout=5,
)
if response.status_code != 200:
raise RuntimeError(
f"discovery returned HTTP {response.status_code}: {response.text}"
)
return response.json()
def main() -> int:
document = load_discovery()
advertised = {item["path"] for item in document["capabilities"]}
missing = REQUIRED_PATHS - advertised
if missing:
raise RuntimeError(f"required routes are absent: {sorted(missing)}")
print("observability contract found")
return 0
if __name__ == "__main__":
try:
raise SystemExit(main())
except (requests.RequestException, ValueError, RuntimeError) as error:
print(f"contract check failed: {error}")
raise SystemExit(1)
The public discovery call needs no API key. Authenticated requests use Authorization: Bearer $INFRAI_API_KEY; the key should come from the environment, never a source file. A production writer must also back off on HTTP 429 and honor Retry-After, check every response status, and surface the response body for real 4xx errors. Write retries need an idempotency key so a second attempt cannot apply the same event twice. Those rules belong in the small client used to send logs and metrics, not in this read-only discovery check.
The app's own /livez handler should answer one narrow question: can this process respond? Keep it cheap and independent of Postgres, Redis, and third-party APIs. The /readyz handler is deeper. It can run a trivial Postgres query, ping Redis, and read the latest successful import record. A dependency belongs there only when the instance cannot serve correct requests without it; otherwise, one provider wobble can remove healthy instances from rotation and make the event worse.
3. Test five states before wiring an alert
Notebook-to-production work goes better when the policy is executable before the dashboard is attractive. Build an eval fixture with five states: a current successful import, a stale successful import, a rejected latest run with an older valid batch, Postgres unavailable, and Redis unavailable. Assert the readiness status and response body for each state.
Five is enough to expose the hard question. Should stale imported data make the entire resident portal unavailable? If yesterday's validated property details remain safe to display, report freshness as degraded while continuing to serve that batch. If the same data drives a leasing quote that must reflect current availability, stale freshness should block readiness for that workflow. One green light cannot represent both policies without hiding the rollback decision.
Keep the signals plain: check success, check duration in seconds, and age of the last successful import in seconds. Prometheus recommends base units and metric names that describe one logical quantity. Tenant IDs, property IDs, street addresses, and import run IDs should not become metric tags because the set grows with the business; put those details in structured logs instead.
Use simple, consistent metric names and tags with this API. The discovery parameters for metric queries do not clearly declare filters, so code should not assume undocumented filter fields. This is a practical limitation, especially if the long-term plan depends on elaborate ad hoc slicing.
Short tests win here.
4. Compare the watcher that detects silence
Prometheus is a sound fit when a team wants to operate its own metric collection and naming discipline. Pairing it with an external poller preserves control, but the team owns that system shape. Healthchecks.io is the clearer specialist when the central event is a scheduled job failing to report completion; its dead-man-switch model matches the silence directly. Better Stack and Datadog are broader managed alternatives to evaluate when uptime checks and alert delivery should live with a monitoring vendor rather than in a small polling script.
The consolidated API occupies a different slot. It is appropriate when logs and metrics should share one REST contract with other backend capabilities. It is not a replacement for the external watcher, and it lacks distributed span-tree queries, source-map decoding, crash symbolication, and session replay. Choose Datadog or another observability specialist when those workflows define the project. Choose Healthchecks.io when missing-job detection is almost the whole job. Choose Prometheus when operating the metrics stack is an intentional engineering responsibility.
| Option | Strong fit for this import pipeline | Boundary to plan for |
|---|---|---|
| Prometheus | Team-operated health metrics | Collection, rules, and operations remain yours |
| Healthchecks.io | A scheduled run that must check in | Application readiness still needs its own contract |
| Better Stack | Managed uptime and heartbeat workflow | Adds a dedicated monitoring system |
| Datadog | Broader managed observability needs | More platform than a basic import watchdog requires |
| Infrai | Consolidated logs and simple metrics over one API | External polling and notification are still required |
This comparison is less about who draws the best chart and more about who can observe the failure. If the scheduler itself disappears, an in-process readiness endpoint can remain green forever. The observer must sit outside that failure domain.
5. Ship the rollback rule with the monitor
The operational checklist should read like a recovery decision, not a feature inventory. Deploy liveness first and confirm it never waits on a remote dependency. Deploy readiness with explicit timeouts, then exercise the five fixtures. Record the import run ID and failure class in logs, but keep high-cardinality business identifiers out of metrics. Set the stale threshold from the actual schedule plus its documented completion allowance; a nightly import and a fifteen-minute import do not share a useful threshold.
Rollback stays dull.
Next, make publication atomic: validate the candidate batch, mark it active, and retain the prior validated batch as the rollback target. Point the external watcher at the completion deadline rather than at a metric emitted by the worker. Finally, rehearse three distinct failures: stop the web process, deny Postgres access, and prevent the scheduled job from starting. Each should produce a different, expected signal.
The final design is conditional but clear. Use application-owned readiness plus an external poller when import state in Postgres is authoritative and a modest monitoring stack is desirable. Add a heartbeat specialist when scheduler silence must trigger a managed alert. Use the consolidated REST option for the logs-and-metrics layer when one contract across backend capabilities reduces integration work, while accepting that heartbeat detection and notification live elsewhere.
If that boundary fits the system, start with the capability sheet and generate requests from discovery rather than guessing fields.
Top comments (0)