Short answer: monitor the public API and the scheduled import as two different promises. Use an external request check for reachability, and emit a completion heartbeat only after an import has produced and committed an acceptable result. For a small fintech SaaS, the least complex useful setup is one endpoint probe, one deadline per import schedule, and a rollback rule that depends on data validity rather than process liveness. Pingdom, UptimeRobot, and Healthchecks can be candidates in that design, but the safer choice is the one that preserves these semantics during a deploy and makes a missed result unambiguous.
This split matters because a healthy web process can keep returning 200 while yesterday's account balances remain in the database. The opposite can happen too: an upstream bank may deliver an empty but valid file while the API remains perfectly available. One green light cannot describe both conditions.
What is better for small SaaS uptime monitoring?
Define health from the user's expected outcome: fresh, validated financial data became visible before its deadline. A scheduler starting is only evidence of intent. A worker exiting is stronger, but it still does not prove that the transaction committed or that the result passed a domain check.
For an import expected every 15 minutes, keep four timestamps: scheduled, started, committed, and heartbeat acknowledged. Also record a low-cardinality outcome such as success, rejected, or failed. Do not put account IDs, file names, or exception messages into metric labels; Prometheus instrumentation guidance warns that every distinct label set creates another time series. Put high-detail context in logs, linked by a run ID.
The important edge case is zero rows. Zero may mean an upstream outage, a holiday, or a legitimate empty interval. Monitoring can't invent that business decision. Encode it in validation: perhaps an empty settlement file is acceptable on a configured market holiday but rejected on an ordinary processing day. Only the accepted branch sends success.
No result, no pulse.
Implementation walkthrough: send success after the commit
The data flow is small. A scheduler creates a run ID, the importer stages rows, validation decides whether the batch is publishable, and a database transaction promotes it. After the commit returns, the worker sends a heartbeat to a pseudonymous URL. Separately, an outside probe calls a shallow health endpoint. The heartbeat detects silence; the probe detects an unreachable application. Neither substitutes for an evaluation of imported data.
Here is a compact Python shape. The standard library keeps the example portable, while the injected functions leave database and domain details where they belong.
from dataclasses import dataclass
from datetime import datetime, timezone
from typing import Callable, Iterable
from urllib.request import Request, urlopen
@dataclass(frozen=True)
class ImportResult:
run_id: str
accepted_rows: int
committed_at: datetime
def send_success_heartbeat(url: str, run_id: str) -> None:
request = Request(
url,
data=b"",
method="POST",
headers={"User-Agent": "scheduled-import/1", "X-Run-ID": run_id},
)
with urlopen(request, timeout=5) as response:
if response.status // 100 != 2:
raise RuntimeError(f"heartbeat rejected with status {response.status}")
def run_import(
run_id: str,
rows: Iterable[dict],
validate: Callable[[list[dict]], None],
commit: Callable[[list[dict]], None],
heartbeat_url: str,
) -> ImportResult:
staged = list(rows)
validate(staged)
commit(staged)
result = ImportResult(
run_id=run_id,
accepted_rows=len(staged),
committed_at=datetime.now(timezone.utc),
)
send_success_heartbeat(heartbeat_url, run_id)
return result
The ordering is deliberate. Sending before commit creates a false success if the transaction later fails. Sending from a finally block is worse because failure and success become indistinguishable. A heartbeat delivery failure after a successful commit should be retried with a stable run ID, but it must not cause the import transaction to run again blindly. That separation protects against duplicate financial records.
This is also where notebook-to-production work often changes shape. In a notebook, “the dataframe has rows” can feel like completion. In production, the meaningful boundary is the durable commit plus domain validation. Turn the same validation cases into an eval harness: normal batch, permitted empty batch, malformed amount, duplicate run ID, late upstream response, and heartbeat timeout. The harness should assert both stored state and emitted outcome.
Decision test: can rollback preserve the missed run?
Comparing Pingdom, UptimeRobot, and Healthchecks by a feature grid starts too late. First write the rollback contract, then test each candidate against it without assuming that similarly named checks behave the same way.
Use a staging job with a 15-minute schedule and deliberately exercise three states. Version A commits and signals success. Version B starts but fails validation, so it must remain silent and cross the deadline. A rollback to A then commits successfully, and the monitor must recover without losing the earlier missed-run event. During that sequence, keep the run IDs visible and ask a reviewer to reconstruct the order from the retained evidence alone. The failed B run must not disappear when A resumes, the delayed alert must refer to the missed deadline rather than the later recovery, and the successful rollback must not replay B's staged rows. If any of those statements can't be demonstrated, a green dashboard is premature. This sequence reveals more than a polished screenshot because it exercises the exact moment when availability and data correctness disagree.
Test the miss.
The decision table is intentionally about observable behavior rather than vendor promises. There is a real trade-off: this two-signal design is a poor fit for a team that can't own import validation or define a completion deadline. In that case, fix the job contract before selecting a monitor; otherwise every candidate will report a precise version of the wrong state. It is also limited when the import is continuously streaming rather than scheduled, because a missed-heartbeat deadline may not represent progress. A lag or watermark metric is the better signal for that workload.
| Test | Required evidence | Reject the setup when |
|---|---|---|
| API process unavailable | External request failure is recorded | A local self-check stays green |
| Import starts but never commits | Deadline expires without success | A start event resets the timer |
| Invalid zero-row batch | Validation failure remains distinct | Any completed process counts as healthy |
| Deploy and rollback overlap | Run IDs preserve event order | The newest arrival overwrites history |
| Heartbeat delivery is retried | One logical run stays identifiable | Retries create separate successes |
Then evaluate operational fit: can the team express the real schedule and grace period, authenticate signals without placing credentials in logs, route alerts to an owned destination, retain enough event history for an incident review, and export or reproduce the configuration? Also check US and EU data-handling requirements against the actual contract and deployment being considered. Region labels alone are not evidence of where monitoring payloads, alert metadata, and logs are processed.
Run this trial for at least several schedule windows, including one controlled failed deployment. That number is not a benchmark; it is the minimum shape of the behavior you need to observe. A monitor that has never seen an intentional miss has not yet demonstrated the core path.
Trade-offs: shallow probes, richer diagnosis
The API health endpoint should answer a narrow question: can this instance serve traffic through its required local dependencies? It should not trigger an import, call every upstream institution, or perform an expensive model inference. Deep dependency fan-out makes a probe noisy and can amplify an outage.
Probe less.
Use metrics for bounded aggregates, logs for run-level evidence, and heartbeat state for deadline detection. Severity needs discipline too. RFC 5424 defines standardized severity levels from Emergency through Debug. Map alert urgency to user impact and required response, rather than labeling every missed execution as the highest severity. A late sandbox import and stale production balances should not page the same way.
Prompt and model costs belong nearby when an AI step classifies transaction descriptions. Record bounded counters for model attempts and accepted outputs, plus token usage at an aggregate level. Never attach raw prompts or customer identifiers as metric labels. The import should also define what happens when that classifier is unavailable: reject the batch, publish a prevalidated subset, or defer enrichment. Pick one behavior and test its rollback path.
The monitor is not the policy engine. It reports that an expected event did or did not arrive; application validation decides whether the event deserves to be called success.
Operational checklist: protect the signal through deployment
Before release, confirm that production and staging use separate check identities, secrets are injected rather than printed, and the grace period covers expected scheduling jitter without hiding a genuinely late batch. Deploy the check definition with reviewable configuration where possible. During rollout, watch the old and new worker versions independently until only one owns the schedule, because overlapping workers can manufacture reassuring pulses while duplicating work.
After a controlled rollback, inspect the full chain: one run ID in scheduler logs, validation outcome, one durable commit, heartbeat acknowledgment, and the external API probe. Re-run the eval cases whenever schedule logic, validation, transaction boundaries, or alert routing changes. Review label cardinality before adding dimensions, and sample the actual alert message so it contains the schedule, environment, last accepted completion, and runbook owner without exposing financial data.
That is the selection criterion: preserve truthful state across failure and rollback. Dashboards, notification channels, and commercial terms matter after the signal can be trusted. A small SaaS does not need many checks to begin; it needs two checks whose meanings stay stable on the worst deployment day.
Further reading
- Prometheus, “Instrumentation”: https://prometheus.io/docs/practices/instrumentation/
- IETF RFC 5424, “The Syslog Protocol”: https://datatracker.ietf.org/doc/html/rfc5424
Top comments (0)