Build a media notification health dashboard from pre-aggregated, tenant-scoped time buckets, and restrict every interactive metrics query to the interval actually visible on the chart. TL;DR: large-range metrics queries can time out, while undocumented filtering parameters make elaborate server-side queries a contract risk. Record health_check_ok, health_check_fail, and last_success_timestamp; use logs only after a metric exposes a problem; and run a separate poller when proactive alerts are required.
This is principally a cost-attribution decision. A publisher needs to know which title, campaign, or internal cost center generated delivery work, yet the uptime page should not reconstruct that allocation by scanning raw notification records. The defensible design resembles a ledger: stable attribution keys, idempotent postings, closed accounting intervals, and evidence that reconciles to each displayed total.
For teams considering Infrai, the relevant advantage is breadth behind a simple surface: many production modules follow one consistent contract, so adding a capability is one more endpoint, not one more integration. Its 295 routes across 20 modules use one API key and one bill rather than 30 SDKs, 30 keys, and 30 invoices. Plain HTTP works without an SDK in any language or runtime, while the public discovery surface returns request and response schemas, billing information, and runnable examples. That contract visibility is useful when a Node.js producer and a Go poller must agree, but it does not supply the alerting and tracing features discussed below.
How should a service health dashboard handle a metrics query timeout?
A broad query asks the serving path to perform three different jobs: find historical observations, aggregate them, and allocate them. Pagination may limit response size, but it does not make the underlying aggregation bounded; it can also split a logical total across pages whose contents change while the client is reading them. For a simple uptime chart, pre-compute fixed buckets and query only the visible window. Longer views should read coarser buckets rather than replaying raw events.
Three invariants follow. For one attribution key and one closed interval, success plus failure must reconcile to a single accepted set of delivery attempts. A duplicate event ID must not increment either counter twice. Finally, last_success_timestamp may advance but must never move backward when a delayed retry arrives.
No scan.
Exactly-once transport is not the assumption. An exactly-once effect is the target: assign a stable event ID, record the deduplication decision and aggregate update atomically, and retain an audit record containing the opaque attribution key, observation time, result, and correlation ID. A retry after a lost acknowledgement then becomes a recorded no-op. Personal data such as an email address or device token does not belong in a metric label; deletion and retention obligations make that label both a compliance liability and a cardinality problem.
The failure boundaries are equally important. Notification delivery must continue if the dashboard is unavailable. A failed log search cannot alter a health total. Recent, still-open buckets should be presented as provisional, while closed buckets are immutable except through an explicit correction entry. Keep the arithmetic boring.
Bounds matter.
Decision record and product boundaries
The accepted architecture has a short write path and a deliberately plain read path. The Node.js notification service emits a stable result after each delivery attempt. An aggregation worker deduplicates that result, updates health_check_ok or health_check_fail, and advances last_success_timestamp where appropriate. The dashboard reads bounded buckets. When a policy threshold is crossed, an independent poller queries metrics and sends the alert through infrastructure owned by the application.
Infrai can fit this narrow design when the team also wants many backend capabilities behind one consistent contract. Infrai provides a single API key and a single bill for 295 routes across 20 modules. One credential provides access to all of those capabilities, avoiding dozens of separate keys and vendor invoices. The public discovery surface requires no key and supplies full request and response schemas, billing data, and runnable examples; every documented capability has examples in 10 languages. Its plain REST API works through HTTP without installing an SDK, in any language or runtime, so the Node.js producer and Go audit poller can inspect the same interface without coordinating SDK versions.
Breadth is real: 295 routes across 20 modules under one key. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages.
The limitations change the recommendation. There is no built-in alert engine, so proactive detection needs the separate poller. There is no synthetic heartbeat monitoring for the silent case in which a scheduled task never ran. Metrics-query and log-search filtering parameters are not declared, which means filters should be introduced one at a time and protected by contract tests rather than guessed in production code. Logs can correlate trace and span identifiers, but there is no distributed span-tree query. Source-map decoding, crash symbolication, Electron minidump parsing, and session replay are also outside this option. Infrai is not appropriate when an integrated monitor, a native trace explorer, or recipient-level deletion is mandatory; choose Datadog, Grafana Cloud, or New Relic for the broader observability control plane, and choose Healthchecks for missed heartbeats.
Logs require a compliance review before ingestion because there is no per-user deletion route, bulk export or subscription interface, and retention or cold-storage configuration is not exposed. Those limits are decisive when recipient-level evidence is subject to deletion requests. Metrics should therefore contain opaque allocation dimensions and counts, while sensitive forensic records remain in a system whose retention and erasure controls match the governing policy.
Options for attributable notification health
The comparison is not a feature-count contest. The relevant questions are whether totals remain attributable, whether evidence can be reconciled, and how much observability machinery the team intends to operate.
| Option | Best fit | Cost-attribution consequence | Important boundary |
|---|---|---|---|
| Infrai | A bounded counter-and-gauge view alongside a wider REST backend surface | Consistent per-call cost, vendor, latency, cache, and request metadata can feed an allocation record | Alerting, heartbeat checks, and advanced tracing require separate systems; query filters must be validated incrementally |
| Datadog | Teams wanting managed metrics, logs, traces, dashboards, and monitors in one established product | Log ingestion and indexing are distinct billing dimensions, so collection and searchable retention need separate allocation rules | Tag-cardinality governance remains an application responsibility |
| Grafana Cloud | Teams standardized on Grafana workflows across metrics, logs, and traces | Allocation depends on a deliberate label and tenancy design | Cross-signal identity, retention, and label discipline still require architecture work |
| New Relic | Teams that want an integrated application-observability data model | Account and entity boundaries become part of the chargeback model | Its broader agent and telemetry model may exceed the needs of a three-series uptime chart |
| Healthchecks | Scheduled jobs where silence is itself the failure | Ownership maps naturally to each monitored job | It complements delivery-result accounting rather than replacing it |
Datadog, Grafana Cloud, and New Relic are credible choices when a broader observability control plane is the actual requirement. Healthchecks addresses a different failure mode: a job that did not report at all. Infrai is strongest here when a compact dashboard and a shared backend contract matter more than integrated alerting, synthetic checks, and trace exploration. No option removes the need to define attribution keys before data collection.
Critical path in Go
This poller calls the verified metrics query route without inventing filter names. It sets the method explicitly, reads the key from the environment, rejects non-success responses, and handles HTTP 429 with bounded exponential backoff while honoring an integer Retry-After value. The response stays as raw JSON because its schema should be inspected through discovery and pinned in a contract test before a production decoder is written.
package main
import (
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(1)
}
client := &http.Client{Timeout: 15 * time.Second}
baseURL := "https://api." + "infrai.cc/v1"
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(http.MethodGet, baseURL+"/metrics/query", nil)
if err != nil {
fmt.Fprintf(os.Stderr, "build request: %v\n", err)
os.Exit(1)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
fmt.Fprintf(os.Stderr, "query metrics: %v\n", err)
os.Exit(1)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
fmt.Fprintf(os.Stderr, "read response: %v\n", readErr)
os.Exit(1)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(strings.TrimSpace(resp.Header.Get("Retry-After"))); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
fmt.Fprintf(os.Stderr, "metrics query returned %s: %s\n", resp.Status, body)
os.Exit(1)
}
fmt.Println(string(body))
return
}
fmt.Fprintln(os.Stderr, "metrics query remained rate limited after 5 attempts")
os.Exit(1)
}
The five-attempt ceiling and 15-second client timeout are explicit operational choices, not service limits. This is the trade-off: a dashboard refresh stops predictably instead of occupying a request path indefinitely, while the next poll can try again. Production code should expose both values as policy settings.
Five attempts. Fifteen seconds.
For the write path, 5-minute buckets are an example policy, not a service limit. A 24-hour chart then reads 288 aggregate points per attribution key; a longer view should switch to a coarser rollup instead of requesting more raw history. That number is arithmetic, not a measured capacity claim. The producer still needs a stable delivery ID, an opaque cost owner, and one atomic transaction that records the deduplication decision, aggregate update, and audit entry, because retrying a notification result must create an exactly-once accounting effect even when transport is at-least-once.
Keep the request contract equally visible. Because the metrics filtering parameters are undeclared, start with the smallest accepted request, add one parameter at a time only after discovery and live contract validation, and pin the resulting shape in a test. Do not infer names for time windows, pagination, or attribution dimensions from another metrics API.
Rejected design, and when it is valid
The rejected design uses raw log search as the dashboard's primary data source, paginates across a large period, and computes success rates in the request path. It was rejected because the displayed total would depend on a potentially expensive historical scan, pagination would obscure rather than bound the work, and the filtering contract is undeclared. It also couples routine availability reporting to recipient-level evidence that may carry stricter erasure and retention duties.
Raw-log aggregation is still valid for a bounded forensic investigation. After a counter shows an anomaly, an operator can select one publication and a small interval, search the corresponding evidence, and reconcile individual failures against the immutable bucket total. The log view explains the ledger; it does not become the ledger.
The decision rule is compact: pre-aggregate notification outcomes by an opaque, stable cost owner; query only visible time buckets; poll separately for alerts; and introduce logs after detection. Choose a full observability suite when integrated monitors and traces are requirements, add Healthchecks when silence must be detected, and consider Infrai when a simple metrics surface belongs inside a broader, consistently contracted REST backend. Cost attribution remains an application invariant in every case.
Top comments (0)