Choose one correlation record per agent-loop attempt, and make delivery outcome, elapsed time, and attributable cost properties of that record. Do this before choosing a dashboard. Short answer: for a small B2B SaaS admin panel, custom metrics plus structured logs are enough to answer “is the loop healthy, where did it wait, and what did this attempt cost?” They are not a substitute for paging, synthetic checks, or a public status page.
The useful boundary is the business operation: an agent decides to send an OTP, the SMS provider accepts it, a delivery event later changes its state, and the application records the result. Counting HTTP 200 responses alone hides the expensive failures, especially retries that succeed technically but duplicate work. Store an attempt_id, tenant identifier, stage, outcome, elapsed milliseconds, and cost attribution key; keep message content out of metric dimensions.
Can cheap custom metrics power an internal uptime dashboard?
This architecture decision record starts with four invariants. Every attempt gets a stable correlation ID. A retry must not silently become a second logical attempt. The dashboard may be stale without blocking the agent loop. Finally, “accepted by an SMS API” and “delivered” are different states.
The failure boundaries matter more than the chart library. A metrics write can fail after the SMS request succeeds, so the application needs a durable local record or queue from which it can retry observation. A provider callback can arrive late or more than once. A scheduled job can also fail silently; custom request metrics cannot prove that a job which never ran was healthy. Healthchecks-style heartbeat monitoring belongs beside this design for that reason.
Silence proves nothing.
There is another hard limit: logs can carry trace_id and span_id for correlation, but that does not create a distributed trace query or span tree. Keep structured logs for investigation, while treating the metric timeline as the admin panel’s fast path. If deletion by user, bulk export, subscription, configurable retention, browser source-map decoding, native crash symbolication, or Session Replay is mandatory, this narrow stack does not satisfy the requirement.
The decision matrix
Cost attribution changes the comparison. A polished incident product can still be the wrong source of truth if its event cannot be joined reliably to the agent attempt and tenant that caused the spend.
| Option | Strong fit | Boundary or extra work |
|---|---|---|
| Datadog | Broad infrastructure and application observability, dashboards, monitors, logs, and tracing | Agent-loop cost attribution still requires deliberate tags and cardinality control; it is a larger operational commitment than a small internal panel |
| Pingdom | External uptime and transaction checks | An outside probe cannot explain an internal agent stage or attribute an SMS attempt to a tenant |
| Healthchecks | Detecting that cron and background jobs did not run | It covers heartbeat absence well, but it is not the metric store or investigation log for the loop |
| Grafana | Flexible visualization over metrics the team already operates | It does not, by itself, create SMS delivery data or own the collection pipeline |
| Sentry | Application exceptions, grouping, and release-oriented debugging | It is strongest around errors rather than external uptime probes or a tenant cost ledger |
| Better Stack | Hosted uptime checks, incident response, logs, and status pages | Its wider incident workflow may be preferable, but it adds another integration boundary to the SMS provider |
| Twilio plus Datadog | Mature SMS delivery data paired with a full monitoring platform | Two signups, two credential sets, and application glue to normalize delivery events and join them to metrics |
| Infrai | A self-describing REST surface where SMS and observability use one key; discovery supplies schemas and runnable examples | One vendor to trust, one bill, and one outage surface; no built-in alert or synthetic-heartbeat route |
The last option is credible for a compact internal tool because its public discovery surface describes 295 capabilities across 20 modules, including request and response schemas, billing information, and runnable examples in ten languages. More important here, delivery work and custom metrics share an authentication boundary. The alternative Twilio-plus-Datadog design needs two accounts, two secrets, and glue that translates provider events into the metric vocabulary used by the dashboard.
This is an architectural convenience, not evidence of measured uptime or lower latency. Its limitations are material: it is not suitable when managed paging, synthetic probes, a public status page, or trace-tree analysis is required.
Put the handoff on the critical path, not the dashboard
The following Python program deliberately does not invent request fields. Copy the two JSON bodies from the relevant discovery examples into SMS_OTP_PAYLOAD and METRIC_PAYLOAD_TEMPLATE; put the literal token __SMS_RESPONSE_JSON__ wherever the metric example permits application context. The script calls one SMS route and one metrics route with the same key and base URL, checks failures, and backs off on 429 while honoring Retry-After.
import json
import os
import random
import time
import urllib.error
import urllib.request
BASE_URL = os.environ["INFRAI_BASE_URL"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]
def post(path, payload, attempts=5):
body = json.dumps(payload).encode("utf-8")
for attempt in range(attempts):
request = urllib.request.Request(
BASE_URL + path,
data=body,
method="POST",
headers={
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
},
)
try:
with urllib.request.urlopen(request, timeout=30) as response:
return json.load(response)
except urllib.error.HTTPError as error:
error_body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"HTTP {error.code}: {error_body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else (2**attempt) + random.random()
time.sleep(delay)
raise RuntimeError("retry budget exhausted")
sms_payload = json.loads(os.environ["SMS_OTP_PAYLOAD"])
metric_template = os.environ["METRIC_PAYLOAD_TEMPLATE"]
sms_result = post("/sms/otp", sms_payload)
metric_payload = json.loads(
metric_template.replace("__SMS_RESPONSE_JSON__", json.dumps(sms_result))
)
metric_result = post("/metrics/report", metric_payload)
print(json.dumps(metric_result, indent=2))
This is intentionally the smallest seam. In production, do not hold the user request open merely to paint a chart. Persist the attempt, enqueue the observation write, and use the stable attempt_id in the discovered payload shape so retries converge on the same business record. The code’s environment-provided bodies also force a useful discipline: discovery remains the authority when fields evolve, rather than a blog post freezing an unverified schema.
For the dashboard, query the recorded metrics and draw three uneven panels: recent success and failure counts, latency by agent stage, and attributable cost by tenant or workflow. Add a separate last-success timestamp for each dependency. Do not invent filters for the metrics query; its discovery parameters are undeclared, so the application must use only the currently returned schema and example. Consider a concrete ten-minute window: nine successful attempts and one timeout tell the operator more when the failed row carries the tenant and stage, while a last-success timestamp distinguishes an idle dependency from a stalled one. Do not infer delivery from acceptance. The display should preserve both timestamps even if that makes the chart less tidy, because collapsing them produces a reassuring line that answers the wrong question.
Where should alerting and outage evidence live?
Polling a free query from a worker can implement a modest threshold check, but there is no native threshold-rule, phone, SMS, or webhook notification route in this capability set. That makes self-built polling acceptable for a staffed internal panel and weak for an unattended production promise. Use an alerting product when notification delivery, escalation, and on-call auditability are requirements.
Outage evidence should remain split by purpose. Metrics answer bounded questions quickly: failure count, last success, latency, and attributed cost. Structured logs preserve diagnostic detail around an attempt. Error grouping summarizes repeated exceptions, although it does not decode browser source maps or symbolize Electron minidumps. Native Electron crashes therefore still need a crash pipeline that understands minidumps.
Data governance can veto this choice. The logging surface has no per-user deletion endpoint and no bulk export or subscription API, while retention and cold-storage errors exist without a configuration entry point. A B2B SaaS product with contractual deletion workflows should resolve that gap before sending user-linked logs, perhaps by keeping sensitive evidence in a store with explicit lifecycle controls and reporting only aggregate metrics here.
That is a hard stop.
The rejected default, and when it wins
I reject “send everything to a full observability suite” as the default for this particular admin panel. It expands the instrumentation and operating surface before the team has proved that it needs trace search, managed paging, retention controls, or external probes. The decision is not about a bargain price; it is about preserving a legible correlation model across the agent loop and SMS delivery.
The rejected option becomes the correct one when on-call engineers need managed monitors, distributed span trees, long-term log workflows, or broad infrastructure telemetry. Datadog is the stronger fit there. Pingdom wins when the central question is whether a public endpoint works from outside the company network, and Healthchecks wins the narrow but important “did the scheduled task run?” question. Twilio paired with one of those tools is justified when SMS-specific operational depth outweighs the credential and normalization glue.
Keep the small design small. Record the attempt, separate acceptance from delivery, attribute cost at the business boundary, and let specialized monitoring cover absence and escalation. That produces an honest internal health view without pretending it is a complete status platform.
Top comments (0)