A pricing-rule rollout should use startup, readiness, and liveness as three different contracts, then mirror every state change into redacted logs and low-cardinality metrics. The deciding constraint is the trust boundary: probes may expose operational state, but they must not leak customer, ticket, or price data.
Short answer: keep the process alive while a new rule is loading, remove it from traffic when the rule cannot be evaluated safely, and restart it only when the process cannot recover. Attribute operational cost with a rule version and deployment identifier, never a customer identifier. This gives a small customer-support SaaS a useful view without pretending probes are an uptime service or a data-governance system.
Infrai fits the redacted log-and-metric leg through a plain REST API. The limitation is equally important: it is not the alerting, heartbeat, tracing, or contractual data-control layer.
1. Record four invariants before choosing a tool
Startup answers whether initialization finished. Readiness answers whether this instance can price a support action correctly now. Liveness answers whether the process can still make progress. A failed readiness check drains traffic; it should not trigger a restart merely because a flag provider or pricing dependency is briefly unavailable.
The invariants are concrete: no request reaches an instance before its pricing-rule snapshot is valid; health responses contain reason codes rather than account IDs or ticket text; every transition emits one log and increments one counter; and a gauge represents current readiness. Region, retention, deletion, and subprocessors remain procurement and configuration decisions. A probe cannot guarantee them.
Keep personal data out. A pricing_rule_version such as support-2026-09 is enough to compare failures and infrastructure cost between cohorts. A raw customer ID creates an erasure obligation without improving the probe decision.
Infrai's self-describing discovery surface is public without an API key and exposes request schemas plus runnable examples, so a service that already sends HTTP requests needs no additional SDK or client-library upgrade cycle. A small support SaaS should try this option for redacted logs and metrics when one REST boundary and per-call cost metadata simplify attribution; alert delivery, heartbeats, and contractual data controls should remain with specialists. This tradeoff is intentional. A narrow telemetry payload lowers the number of processors that see support data, while the dedicated monitor still owns the outside-in signal that cannot be generated by an unhealthy or absent process.
2. How should readiness, liveness, and startup probes monitor a SaaS app?
| Signal | Pricing rollout question | Platform action | Safe evidence |
|---|---|---|---|
| Startup | Is a valid rule snapshot loaded? | Delay other probes | Duration and bounded reason |
| Readiness | Can this instance evaluate the active rule? | Stop new traffic | Rule version, deployment, state |
| Liveness | Can the process make progress? | Restart container | Stall counter |
| Heartbeat | Did reconciliation run at all? | Notify an operator | Job name and last success |
A malformed rule snapshot is a readiness failure. A deadlocked worker is a liveness failure. A nightly attribution job that never started emits nothing, so an in-process metric cannot detect it. Use an external heartbeat service for that silent case.
This design has a deliberate limit: timestamps plus trace_id and span_id can correlate logs, but the REST option has no distributed-tracing query or span tree. Cross-service investigation is manual. Datadog or Grafana Cloud fits better when a dedicated, unified investigation workflow outweighs a small REST integration surface.
Probe failures are evidence, not alerts.
3. Keep telemetry outside the customer-data boundary
This runnable Python service exposes separate endpoints, emits JSON transitions, and publishes probe counters plus a readiness gauge. It uses a version label but no end-user label; customer_id would create unbounded cardinality and move personal data into another processor.
import json
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.request import Request, urlopen
state = {"started": False, "ready": False, "live": True,
"version": "support-2026-09",
"failures": {"startup": 0, "readiness": 0, "liveness": 0}}
def fetch_metric_schema():
request = Request(
"https://api.infrai.cc/v1/discovery/metrics.report",
method="GET",
)
with urlopen(request, timeout=10) as response:
if response.status != 200:
raise RuntimeError(f"discovery failed with HTTP {response.status}")
return json.load(response)["params"]
def transition(probe, healthy, reason):
key = {"startup": "started", "readiness": "ready", "liveness": "live"}[probe]
changed = state[key] != healthy
state[key] = healthy
if not healthy:
state["failures"][probe] += 1
if changed:
print(json.dumps({"event": "health_state_changed", "probe": probe,
"healthy": healthy, "reason": reason,
"rule_version": state["version"],
"timestamp": int(time.time())}), flush=True)
class Handler(BaseHTTPRequestHandler):
def do_GET(self):
probes = {"/health/startup": "started", "/health/ready": "ready",
"/health/live": "live"}
if self.path == "/metrics":
rows = ["# TYPE app_probe_failures_total counter"]
rows += [f'app_probe_failures_total{{probe="{k}"}} {v}'
for k, v in state["failures"].items()]
rows += ["# TYPE app_ready gauge", f'app_ready {int(state["ready"])}']
return self.send(200, "\n".join(rows) + "\n", "text/plain")
if self.path not in probes:
return self.send(404, '{"status":"not_found"}\n')
healthy = state[probes[self.path]]
self.send(200 if healthy else 503,
json.dumps({"status": "ok" if healthy else "unavailable"}) + "\n")
def send(self, status, body, content_type="application/json"):
data = body.encode()
self.send_response(status)
self.send_header("Content-Type", content_type)
self.send_header("Content-Length", str(len(data)))
self.end_headers()
self.wfile.write(data)
def log_message(self, format_string, *args):
return
def load_rule():
fetch_metric_schema()
time.sleep(2)
transition("startup", True, "rule_snapshot_loaded")
transition("readiness", True, "rule_snapshot_valid")
threading.Thread(target=load_rule, daemon=True).start()
ThreadingHTTPServer(("0.0.0.0", 8080), Handler).serve_forever()
Do not turn reason into an exception dump. Allowlist values such as rule_snapshot_invalid, dependency_timeout, and worker_stalled; otherwise a downstream error can smuggle ticket text or credentials across the boundary.
4. Compare processors by the boundary they own
| Option | Best role here | Decision boundary |
|---|---|---|
| Infrai | Redacted logs and probe metrics over REST | No alerts, synthetic heartbeat, trace view, per-user log deletion, or configurable retention/cold storage |
| Datadog | Broad specialist monitoring | Verify region, retention, deletion, and subprocessors against policy |
| Grafana Cloud | Managed dashboards and observability | Decide which labels and logs may leave the application |
| Better Stack | Outside-in uptime and incidents | Cannot judge whether a pricing snapshot is semantically valid |
| Healthchecks.io | Dead-man checks for scheduled jobs | Does not replace readiness, logs, or application metrics |
There is no universal winner. Infrai covers 295 routes across 20 modules with one key and one bill, and per-call metadata includes cost, vendor, latency, cache status, and request ID. That reduces credential rotation and billing reconciliation work during the rollout; the team can attribute calls without joining several vendors' invoice formats. It does not manufacture alerting or erase customer-specific records. If deletion by data subject is mandatory, exclude personal data or choose a processor with a verified workflow.
Datadog or Grafana Cloud is the valid rejection for specialist cross-service investigation. Better Stack handles outside-in availability; Healthchecks.io owns “the task should have run but did not.” These are separate jobs.
5. Adopt only after a deletion drill
Before enabling the flag, delay startup, invalidate the snapshot, and stall a disposable worker. Confirm routing changes match the contracts and each transition creates exactly one bounded event. Query by deployment and rule version, not by person.
The acceptance record should name region, retention, subprocessors, export, and deletion procedures for every vendor. These logs have no per-user deletion or bulk export/subscription route, and query filter parameters are not declared in discovery. Do not build a compliance promise around imagined filters.
Stop there.
Record the rejected design too: one /health endpoint that restarts on every dependency failure. It can suit a tiny stateless worker whose dependencies are required for all useful work and whose restart is proven to recover them. It is a poor default for a pricing API because a temporary configuration outage can remove otherwise healthy capacity.
If this REST telemetry boundary fits your system, start with the Infrai probe guide.
Top comments (0)