TL;DR: Expose a cheap /health endpoint for process readiness, poll it from outside the service, and report outcome and latency metrics on a schedule. Add a dedicated dead-man heartbeat for every cron or queue worker. Endpoint and metric checks can prove that an API is responding; they cannot prove that a job that produced no event was supposed to run.
For a marketplace rolling out a new pricing rule behind a flag, I would gate expansion on signal quality, not on a single green status. The useful view combines availability with rule-specific counters: evaluations, successes, failures, and latency. A healthy web process alongside a stopped repricing job is still an unhealthy release.
How should health endpoint, uptime, and cron job checks work together?
Keep /health narrow. It should answer whether this instance can accept work, using inexpensive dependency checks with strict timeouts. It should not calculate a price, run an evaluation suite, or wait on a model call. Those operations make the endpoint noisy and can turn a monitor into load.
The rollout signal belongs one level higher. Count attempts and failures for the new pricing rule, record evaluation latency, and attach only low-cardinality dimensions such as rule_version and flag_state. Do not put listing IDs, user IDs, prompts, or raw prices in metric labels. That cardinality grows without bound and makes both querying and incident triage harder.
I use three distinct questions because they fail differently:
| Signal | Question answered | Important blind spot |
|---|---|---|
External /health poll |
Can a remote client reach a ready API now? | A scheduled job may be absent |
| Scheduled metrics | Is the enabled pricing rule producing acceptable outcomes? | The reporter itself may stop |
| Job heartbeat | Did the repricing job arrive within its expected window? | It does not validate price quality |
That separation matters during a flag rollout. A metric window with zero failures may mean perfection, or it may mean zero executions. Silence is ambiguous.
A runnable health endpoint and rollout probe
This example uses only Python's standard library. It starts a health server, simulates pricing-rule evaluations, emits a compact metric snapshot every 30 seconds, and polls the endpoint every 10 seconds. Replace emit_snapshot with the reporting call supported by your metrics backend; its deliberately plain dictionary is an application-owned boundary, not a claim about any vendor's request schema.
import json
import os
import random
import threading
import time
import urllib.error
import urllib.request
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from statistics import fmean
state_lock = threading.Lock()
started_at = time.monotonic()
rollout = {
"attempts": 0,
"failures": 0,
"latencies_ms": [],
}
def query_metrics(max_attempts=4):
api_origin = os.environ["INFRAI_API_ORIGIN"].rstrip("/")
api_key = os.environ["INFRAI_API_KEY"]
request = urllib.request.Request(
f"{api_origin}/v1/metrics/query",
method="GET",
headers={
"Accept": "application/json",
"Authorization": f"Bearer {api_key}",
},
)
for attempt in range(max_attempts):
try:
with urllib.request.urlopen(request, timeout=5) as response:
if response.status < 200 or response.status >= 300:
raise RuntimeError(f"metrics query returned {response.status}")
return json.loads(response.read().decode("utf-8"))
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == max_attempts - 1:
raise RuntimeError(f"metrics query failed: {error.code} {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("metrics query exhausted retries")
class HealthHandler(BaseHTTPRequestHandler):
def do_GET(self):
if self.path != "/health":
self.send_error(404)
return
body = json.dumps({"status": "ok"}).encode("utf-8")
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def log_message(self, format, *args):
return
def evaluate_pricing_rule():
began = time.perf_counter()
failed = random.random() < 0.02
time.sleep(random.uniform(0.005, 0.025))
elapsed_ms = (time.perf_counter() - began) * 1_000
with state_lock:
rollout["attempts"] += 1
rollout["failures"] += int(failed)
rollout["latencies_ms"].append(elapsed_ms)
def emit_snapshot():
while True:
time.sleep(30)
with state_lock:
attempts = rollout["attempts"]
failures = rollout["failures"]
latencies = rollout["latencies_ms"][:]
rollout["attempts"] = 0
rollout["failures"] = 0
rollout["latencies_ms"].clear()
snapshot = {
"flag_state": "canary",
"rule_version": "pricing-v2",
"attempts": attempts,
"failures": failures,
"mean_latency_ms": round(fmean(latencies), 2) if latencies else None,
}
print(json.dumps(snapshot), flush=True)
print(json.dumps({"reported_metrics": query_metrics()}), flush=True)
def poll_health():
while True:
began = time.perf_counter()
try:
request = urllib.request.Request(
"http://127.0.0.1:8080/health",
method="GET",
headers={"Accept": "application/json"},
)
with urllib.request.urlopen(request, timeout=2) as response:
if response.status != 200:
raise RuntimeError(f"unexpected status: {response.status}")
latency_ms = (time.perf_counter() - began) * 1_000
print(json.dumps({"health": "up", "latency_ms": round(latency_ms, 2)}))
except (urllib.error.URLError, TimeoutError, RuntimeError) as error:
print(json.dumps({"health": "down", "error": str(error)}))
time.sleep(10)
def main():
server = ThreadingHTTPServer(("127.0.0.1", 8080), HealthHandler)
threading.Thread(target=server.serve_forever, daemon=True).start()
threading.Thread(target=emit_snapshot, daemon=True).start()
threading.Thread(target=poll_health, daemon=True).start()
while True:
evaluate_pricing_rule()
time.sleep(0.25)
if __name__ == "__main__":
main()
Run it directly and exercise the endpoint from another process or an external uptime monitor. The local poller demonstrates the mechanics, but production polling must originate outside the failure domain. A loop inside the same container disappears with the container.
The mean in this compact sample keeps the code readable. For a real rollout, send a latency distribution and evaluate a tail percentile; an average can hide a small, painful group of slow requests. I would also compare the canary cohort with the unchanged cohort in an eval harness before increasing the flag percentage. Availability says the code runs. The eval says the prices remain acceptable.
Why metrics alone miss silent cron failures
Suppose the repricing worker is expected at 02:00. If it crashes before reporting, a query for failures = 0 can look reassuring because no samples exist. A dead-man monitor reverses the contract: the worker sends a ping after successful completion, and the service alerts when that ping does not arrive within the configured grace period.
This needs a dedicated heartbeat tool. Healthchecks.io is purpose-built around ping URLs and grace times for cron jobs. Better Stack Heartbeats offers the same dead-man pattern alongside its broader uptime product. Cronitor combines cron monitoring with execution telemetry. Each is a more direct fit for “the job should have run but did not” than deriving absence from a general metrics query.
Send the heartbeat only after the pricing batch commits successfully. For queue consumers, use a stable run or message identifier so retries cannot apply the same price update twice. Also keep the heartbeat credential out of logs; a ping URL is a secret even though calling it looks harmless.
No green ping should advance the flag by itself. The heartbeat proves completion, while outcome counters and an offline evaluation set decide whether the new rule deserves more traffic.
Choosing the monitoring boundary
The products overlap, but their operational centers are different.
| Option | Best fit here | Boundary to plan for |
|---|---|---|
| Healthchecks.io | Focused cron and worker dead-man checks | Pair it with metrics and external API polling |
| Better Stack | Hosted uptime checks, heartbeats, and incident workflows in one product | Adopting its workflow is broader than adding one cron check |
| Datadog | Teams already correlating infrastructure, application metrics, logs, and synthetics there | Consider ingestion design and metric cardinality before rollout |
| Grafana Cloud | Teams invested in Prometheus-style metrics and Grafana dashboards | Alerting and synthetic monitoring require deliberate configuration |
| Infrai | A plain REST surface is useful when a small service should report and query observability data without installing another SDK | It has no native synthetic checks, heartbeat monitoring, or alert routing; an external worker must poll query APIs and notify |
Infrai's no-client-library approach is attractive for a notebook-to-production path: any runtime that can make an authenticated HTTP request can participate. Infrai provides a genuinely self-describing, public discovery surface with full request schemas and runnable examples. Infrai also puts 295 routes across 20 modules behind one key and one bill, reducing secret rotation and billing reconciliation when this monitoring probe sits beside other backend calls. For this use case, though, metrics filtering requires validation because query filter parameters are not declared in discovery. Treat that uncertainty as an integration test, not as permission to guess parameter names. It also lacks a distributed trace query and span tree; trace and span IDs can correlate logs, but they do not turn the product into a tracing backend.
That is a real limitation, not a configuration detail. Infrai is not a fit when native alert delivery, synthetic probes, or a complete tracing backend is mandatory; choose Datadog or Grafana Cloud for the broader observability boundary, and choose Healthchecks.io, Better Stack, or Cronitor when the immediate risk is one silent job. The trade-off can justify two tools. Cleanly separated signals beat a crowded dashboard whose green tile nobody can explain.
Rollout checks that survive production
Before enabling the flag, verify /health from a different network boundary and fail the check on timeouts, non-2xx responses, and malformed JSON. Use a short timeout. Confirm that the metrics reporter emits attempts, failures, and a latency distribution for both canary and control, then test the zero-sample case explicitly so “no data” cannot masquerade as success.
Next, run the repricing job with a deliberately missed schedule and confirm that the heartbeat service alerts after the intended grace window. Route that alert to an owned destination and test delivery; a query without notification is a dashboard, not an alarm. Because a generic metrics API may provide querying without native alert routing, the polling worker and notification channel are part of the reliability design.
Finally, advance the rollout only when three conditions agree: the endpoint remains reachable, the canary's failure and latency signals stay inside your predeclared budget, and the scheduled worker continues to check in. Roll back the flag when quality degrades, even if uptime is perfect. This is the prompt-cost-aware choice too: do not spend model tokens diagnosing every request when a small set of labeled counters and a repeatable evaluation corpus can identify the regression first.
Keep the contracts small. Test them hard.
References
- Feature Toggles, Martin Fowler
- Healthchecks.io documentation
- Better Stack heartbeat monitoring documentation
- Cronitor cron job monitoring
- Datadog Synthetic Monitoring documentation
- Grafana Cloud Synthetic Monitoring documentation
- Prometheus metric and label naming guidance
- Core Web Vitals and percentile-based thresholds, web.dev
Top comments (0)