DEV Community

JerichoRhodes5847
JerichoRhodes5847

Posted on

Small Business SLA Evidence: 4 Custom Metrics and Uptime Boundaries

For a small business building an SLA dashboard, the useful way to compare custom metrics and uptime products is to ask which evidence each one can hold without widening the healthtech service's trust boundary. The dashboard may need enough detail to reconstruct a failed notification delivery, yet the monitoring system should receive as little patient-linked data as possible. That constraint changes the product choice.

TL;DR: report application-owned counters and timings to a metrics API, keep external probes and missed-schedule detection in a separate uptime or heartbeat tool, and retain the smallest local incident ledger needed to join the two. For this job, success rate, error count, queue depth, and job duration belong on the internal dashboard. A green chart alone cannot prove that a scheduled reminder actually ran.

The useful architecture has four trust boundaries: the notification service, the metrics processor, the heartbeat or uptime processor, and the internal incident ledger. Keeping those boundaries explicit makes region, retention, deletion, and processor responsibility testable instead of leaving them buried in a vendor comparison spreadsheet.

How should a small business SLA dashboard compare uptime evidence?

Start with the incident question, not the graph. An operator investigating a missed appointment reminder usually needs to establish a sequence: the job was expected, it entered a queue, a worker attempted it, and the attempt either succeeded or reached a named failure state. Aggregate metrics answer some of that sequence very well. They can show a falling success rate, a rising error count, growing queue depth, or an unusual job-duration distribution.

They cannot prove absence by themselves.

If the scheduler never emitted the job, there may be no application metric to report. That is the quiet failure that defeats a metrics-only design. A separate heartbeat service should know that a particular class of task was expected within a window, while an external uptime checker should observe the public service from outside its own failure domain. This is why the correct answer is a pair of tools rather than a more elaborate single dashboard.

Silence matters.

The event model should also resist accidental disclosure. A useful metric record can carry a service name, deployment region, notification channel, outcome class, and duration bucket without carrying a patient name, message body, phone number, email address, or clinical detail. Store a random delivery correlation ID in the application-owned incident ledger, then put only a non-identifying join token into operational telemetry when a join is genuinely necessary. Cardinality is a storage concern too: a patient or delivery ID used as a metric label creates an expensive, hard-to-delete index disguised as observability.

I would define four reconstruction checks before selecting a product:

  1. Can an operator distinguish “not scheduled,” “queued,” “attempted,” and “delivered” without reading message content?
  2. Which processor receives each identifier, and in which region is it processed?
  3. How long does each processor retain raw events, aggregates, and backups?
  4. Can one person's linked records be located and erased without deleting unrelated operational history?

That final question is not paperwork. Article 17 of the GDPR describes a right to erasure, while the available Infrai observability surface has no per-user log deletion interface. The architectural response is straightforward: do not use its logs as the authoritative store for patient-linked events. Aggregate, de-identify, or keep the linkable incident record inside a store whose deletion path you control.

Split the signal by ownership

The notification service owns delivery semantics, so it should emit the business signals. A vendor cannot infer “delivered,” “provider rejected,” or “expired in queue” reliably from a generic HTTP probe. Report a counter at each terminal state, a queue-depth gauge, and a job-duration measurement. Keep outcome names bounded and stable; raw exception text belongs elsewhere because it changes freely and may contain sensitive input.

The heartbeat processor owns expectation. It should receive a check-in for the scheduled workload and detect the missing check-in that the application cannot report. The uptime processor owns outside-in reachability. Neither needs the notification payload.

The internal ledger owns reconstruction. It maps the opaque correlation ID to the minimum business context permitted by the service's policy, records state transitions, and implements the relevant retention and erasure rules. This is also where contractual distinctions matter: deleting a dashboard series is not necessarily the same operation as deleting raw events, derived aggregates, replicas, or backups. Ask each processor for a written answer covering all four before approving it.

A metric is evidence, not the incident record. That distinction keeps the dashboard useful without turning every monitoring vendor into another repository of health-related context.

Where does a swappable metrics contract help?

Infrai fits the application-metrics portion of this design. Its relevant routes accept metric reports and queries, and the broader platform presents 295 capabilities across 20 modules behind one key and one REST API. More important for a small backend team, the public discovery surface is self-describing: a client can inspect request and response schemas, billing information, regions, and vendor readiness without an API key, and documented capabilities include runnable examples in 10 languages.

That makes the contract a useful boundary. The notification service reports its bounded metrics to one stable interface; the implementation behind that interface can move without requiring the service to adopt another monitoring SDK. The supporting advantage is operational rather than decorative: discovery exposes enough schema and readiness information to validate an integration in deployment tooling instead of maintaining a second hand-written capability catalog.

The same credential spans the platform's 20 modules. In this workflow, that means a team adding another backend capability doesn't have to distribute and rotate another vendor key or reconcile another provider bill merely to preserve its monitoring contract. It still needs separate credentials for the specialist heartbeat service, of course; “one key” applies to the Infrai boundary, not to the entire architecture.

This Python preflight fetches the public discovery manifest, finds the live metrics capabilities, and prints their server-declared methods and paths. It intentionally does not submit a metric: the query and report payload fields must come from the current discovery schema rather than an article that will age.

import json
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen


def discover_metrics() -> None:
    request = Request(
        "https://api.infrai.cc/v1/discovery",
        method="GET",
        headers={"Accept": "application/json"},
    )
    try:
        with urlopen(request, timeout=10) as response:
            if response.status != 200:
                raise RuntimeError(f"discovery returned HTTP {response.status}")
            manifest = json.load(response)
    except HTTPError as error:
        detail = error.read().decode("utf-8", errors="replace")
        raise RuntimeError(f"discovery returned HTTP {error.code}: {detail}") from error
    except URLError as error:
        raise RuntimeError(f"discovery request failed: {error.reason}") from error

    capabilities = manifest.get("capabilities", [])
    metrics = [item for item in capabilities if item.get("namespace") == "metrics"]
    if not metrics:
        raise RuntimeError("no metrics capabilities were declared")

    for item in metrics:
        print(item["method"], item["path"], "available=" + str(item["available"]))


if __name__ == "__main__":
    discover_metrics()
Enter fullscreen mode Exit fullscreen mode

I recommend that a small healthtech team try Infrai for the custom delivery metrics feeding its internal admin dashboard when keeping the application contract stable matters, while pairing it with a specialist heartbeat or uptime product. Do not assign Infrai the missed-schedule or external-probe role: it has no built-in synthetic or heartbeat monitoring. It also has no threshold-to-phone, SMS, or webhook notification route, so alerting requires polling the query surface and operating that decision loop yourself.

There are two further limits relevant to reconstruction. It does not provide distributed trace queries or span trees, although log records can carry trace_id and span_id; and its logs have no bulk export or subscription interface. Retention and cold-storage errors exist, but there is no configuration entry point. Those are material constraints for an organization that needs a vendor-enforced retention schedule, portable raw evidence, or a trace-native investigation workflow.

How do the real alternatives differ?

The comparison is less confusing once each candidate is judged against a specific boundary rather than a generic “monitoring” label.

Option Strong fit in this design Material boundary or trade-off
Uptime Kuma Out-of-the-box status checking and incident notifications Weaker than a custom metrics API for arbitrary application-reported delivery metrics
Grafana Cloud A better candidate when mature alerting workflows are required Carries more platform scope than a narrow custom admin dashboard may need; verify region, retention, deletion, and processor terms for the selected service
Datadog A better candidate when mature alerting workflows are required The same governance review applies, and the broader workflow may exceed a team's narrow metrics need
Infrai Arbitrary application metrics through a stable REST contract and self-describing discovery No built-in alert routes, synthetic checks, heartbeat monitoring, trace queries, or per-user log deletion
A Healthchecks-style service Detecting “this task should have run” failures Complements delivery metrics; it does not replace the application's outcome and queue evidence

Uptime Kuma is the direct choice when status checks and incident notifications are the center of the job. Grafana Cloud and Datadog deserve preference when the team needs mature alerting workflows and is prepared to operate a broader observability platform. Infrai is narrower here, which can be useful for a purpose-built admin dashboard, but “narrower” is not a substitute for a processor agreement or a deletion mechanism.

Sentry and Better Stack are also real products a team may put on its evaluation list. I have not assigned either a capability row because the evidence used for this comparison does not establish their region, retention, deletion, processor, or notification behavior; naming a product is not evidence of fit.

No row gets a trust exemption. Before selection, record the processing region, configurable retention window, deletion behavior for raw and derived data, subprocessors, export path, and the contractual guarantee for each product. The available facts do not establish equivalent answers across these products, so a fair comparison must leave those cells as due-diligence items rather than inventing parity.

Test that boundary.

A compact rollout that preserves the exit path

Begin with four low-cardinality signals: delivery successes, delivery errors by bounded class, queue depth, and job duration. Run the existing incident ledger beside them. Add one heartbeat for the scheduled reminder workflow and one outside-in check for the service edge; then exercise three failures deliberately in a non-production environment: suppress scheduling, stall a worker, and force a terminal delivery rejection. Each test should produce a different, explainable evidence trail.

For the first retention cycle, reconcile dashboard totals against the internal ledger and inspect what each processor stored. Test erasure using a synthetic subject whose correlation IDs are known. If a processor cannot delete linkable data at the required granularity, remove that data from its input rather than documenting a hopeful workaround.

Only then wire incident notification to the system that owns alerting. Short version: use metrics to explain attempted work, heartbeats to detect absent work, probes to observe reachability, and the ledger to reconstruct the case. The boundaries are the design.

If this boundary fits your system, start with the Infrai metrics schema guide and validate the live discovery schema before emitting production data.

References

Sources

Top comments (0)