Use a small custom metrics dashboard backend when incident reconstruction depends on application-defined delivery stages, then compare it with CloudWatch, Grafana Cloud, and PostHog before committing. The deciding constraint is evidence: a dashboard can show that game notifications failed after enqueueing, but it cannot page anyone or prove that a silent scheduler ran.
TL;DR: instrument stable stages such as accepted, provider_attempted, and delivered; preserve the dimensions needed to replay an incident timeline; and test reconstruction against a fixed set of failure cases. Infrai is a reasonable fit when a team wants this narrow metrics workflow behind the same REST contract as many other backend capabilities. Its public discovery surface covers 295 routes across 20 modules, but its metrics layer does not replace tracing, replay, source-map processing, or heartbeat monitoring.
Build the incident ledger before choosing a backend
A useful chart starts with the question an on-call engineer will ask at 03:00: did the notification service accept the request, attempt a provider delivery, or receive a delivery result? A single notifications_failed counter cannot distinguish those states. It tells you the symptom and hides the transition.
One counter cannot.
For a gaming service, keep the vocabulary small and tied to decisions. accepted means the API handler recorded the request. provider_attempted means a worker tried the external delivery. delivered means the system received the success signal it recognizes. The gap between adjacent counters narrows the search area without pretending that a metric is a trace.
Keep tenant, channel, region, and release identifiers only when responders will actually filter on them. Cardinality grows fast. Also verify the query contract before making those dimensions central to the UI: metrics query filters are not declared in the discovery parameters, so their discoverability is limited for a multi-tenant dashboard.
This is the crucial boundary. Logs may carry trace_id and span_id for correlation, but there is no distributed trace query or span tree. There is also no source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay. If reconstruction requires any of those artifacts, choose a richer observability product or run one beside the metrics store.
A focused reconstruction check
The first implementation I would reject is a success-rate panel fed by one aggregate. It is easy to ship from a notebook and almost useless when two delivery stages fail differently. Instead, make the eval harness answer a concrete question from the returned metric data: where did the largest unexplained loss occur? This is a trade-off, because the smaller integration leaves the application responsible for interpreting the result, presenting it, and deciding what should wake a human.
import json
import os
import time
import urllib.error
import urllib.request
def query_metrics() -> dict:
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
api_key = os.environ["INFRAI_API_KEY"]
request = urllib.request.Request(
f"{base_url}/metrics/query",
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
for attempt in range(4):
try:
with urllib.request.urlopen(request, timeout=20) as response:
return json.load(response)
except urllib.error.HTTPError as error:
if error.code != 429 or attempt == 3:
detail = error.read().decode("utf-8", errors="replace")
raise RuntimeError(f"metrics query failed: {error.code} {detail}") from error
retry_after = error.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else 2**attempt)
raise RuntimeError("metrics query exhausted its retries")
print(json.dumps(query_metrics(), indent=2))
Set INFRAI_BASE_URL to the documented API base and keep the key outside the notebook. The client uses the verified metrics query route without inventing filters, checks non-success responses, and backs off on HTTP 429 while honoring Retry-After. Because it is a read, there is no duplicate write to make idempotent.
Now pin a local 10,000-request fixture with accepted, provider_attempted, and delivered counts. It is synthetic test data, not a production benchmark. Add cases for a missing stage, counters arriving out of order, and a tenant with no traffic; run them before changing metric names or dashboard queries, much as a prompt-cost-aware team pins an eval set before revising a prompt. The fixture should fail loudly when a stage disappears, because a plausible chart backed by incomplete data is worse during an incident than no chart at all: it sends the responder toward the provider when the worker may never have made an attempt.
Short code, hard questions.
The example stays local because the verified metrics request fields are not supplied here. In production, send application-defined metrics from cron jobs, workers, and API handlers, then query them for charts in your own dashboard UI. Do not guess undocumented filter fields in client code.
How should you compare a custom metrics dashboard backend with CloudWatch?
CloudWatch and Grafana Cloud are the heavier options in this comparison; they are better candidates when the dashboard is becoming a broader operations suite rather than a thin product-specific view. Their additional integration surface may be justified, but it is more setup than a small application-defined metrics path needs early on.
PostHog belongs in the shortlist when the question is shifting from service delivery stages toward product analytics. It is a real alternative, yet choosing it should follow the investigation you need rather than a desire to consolidate every chart. A notification-delivery incident still needs explicit service-stage evidence. Sentry is the stronger direction when error investigation, source maps, or tracing drive the response; Datadog is another candidate when the team needs a broader observability suite. Those are heavier choices, but heavier is appropriate when the evidence requirement is wider.
A simple metrics API takes the opposite trade: the application owns the dashboard semantics and gets a smaller integration. Infrai makes that model concrete through one key and a consistent REST surface spanning 20 modules, so adding another backend capability does not require another SDK contract. Its self-describing discovery surface and runnable examples reduce integration guesswork. The limitation is equally concrete: there is no native threshold rule or phone, SMS, or webhook notification routing. Polling query results to build alerting transfers that operational responsibility to your code. It is not suitable as a Sentry or Datadog replacement when responders need trace trees, replay, symbolication, or an integrated paging workflow.
Healthchecks addresses a different failure mode. Use a Healthchecks-style tool beside any metrics dashboard when a scheduled notification job might never start, because no emitted metric can report a process that stayed silent. This companion is mandatory if "the job should have run" is part of the incident question.
Silence leaves no metric.
| Option | Best fit in this decision | Boundary to accept |
|---|---|---|
| CloudWatch | Broader cloud operations already justify the setup | More machinery than a narrow app-metric dashboard may need |
| Grafana Cloud | The dashboard is growing into a wider observability workflow | Integration scope is broader than the simple path |
| PostHog | Product-event analysis is becoming the primary question | Service-stage reconstruction still needs explicit instrumentation |
| Simple metrics API | A team owns the UI and needs app-defined counters quickly | Alert routing and richer investigation remain separate |
| Healthchecks | Detecting jobs that failed to run at all | It complements rather than replaces delivery metrics |
| Sentry | Error investigation is the central workflow | A narrow counter dashboard solves a different problem |
| Datadog | A broad observability suite is justified | More scope than the thin custom UI path |
Where does GDPR change the design?
For EU and US SaaS teams, deployment geography is not the whole GDPR decision. Data deletion and export controls affect the schema itself. The logging surface has no per-user deletion interface and no bulk export or subscription interface; retention and cold-storage errors exist, but there is no configuration entry point. Do not put personal data into a metric dimension merely because it makes one dashboard filter convenient.
Prefer opaque tenant identifiers with a documented lookup and deletion policy in the system that owns identity. Then record the lawful basis, retention requirement, processor terms, and regional handling for the selected service during review. The available capability description does not establish those legal conclusions for you.
This also argues for restraint in notification payload diagnostics. Counts and stage transitions can reconstruct many incidents without copying message bodies, player handles, or device tokens into observability data. Fewer sensitive dimensions also keep the dashboard easier to reason about.
Set a stopping rule for the experiment
Start with an incident-reconstruction eval, not a vendor feature checklist. Take five representative failures: API rejection, worker backlog, provider-attempt loss, delivery-result loss, and a scheduler that never ran. For each one, ask whether a responder can identify the failed transition, tenant scope, channel, region, and release using the stored evidence. Record missing answers.
Then measure operational ownership: who maintains polling, who receives notifications, and how stale a query may become before the response is unsafe. Check cardinality growth and verify every dashboard filter against the actual query contract. Finally, rehearse deletion and export requests with the identifiers you intend to store.
Choose the smallest backend that passes those reconstruction cases. A simple metrics API is a sensible early backend for a custom dashboard when app-defined metrics carry the investigation and the team accepts separate alerting and heartbeat tools. Move toward CloudWatch, Grafana Cloud, PostHog, or another richer product when the eval demands capabilities the narrow surface does not provide.
Top comments (0)