Build the smallest internal uptime dashboard by polling recent health metrics and structured logs, aggregating them into green, yellow, or red states, and linking failures to error groups for the same service. Keep the status calculation in your application so a deployment can roll back without migrating dashboard state.
TL;DR: Poll directly rather than designing around a log stream. Store a compact metric for every check, retain detailed logs only around failures, and make the UI state a reproducible function of a recent window. This gives an edtech operations team current email, SMS, and OTP delivery visibility without pretending that a lightweight dashboard is a compliance archive or a complete alerting system.
Start with the storage bill
The recurring bill has three parts: check ingestion, retained data, and dashboard queries. Before choosing a product, quantify the first two with your own payload sizes and retention window. Query pricing and indexing rules vary, so they belong in the vendor worksheet rather than in a universal estimate.
Consider an explicit planning case, not a benchmark: 20 notification services, one check every 60 seconds, and a 1 KB structured log for every result. That produces 28,800 checks and about 28.8 MB of raw logs per day, before indexing or storage overhead. Thirty days keeps roughly 864 MB. Change the interval to 10 seconds and the same design creates six times as many records.
The interval wins.
The useful change is to keep a small status metric for every check but write the detailed health log on a transition, a failure, and a limited sample of successes. You preserve the timeline needed for green/yellow/red calculation while avoiding six copies of an uneventful ok payload every minute. Keep error groups beside the failure logs so an operator can move from “OTP delivery is red” to recent exceptions affecting that service.
There is a cost to this restraint. After the detailed-log window expires, you can recover the fact that a service was unhealthy, but not every response body or diagnostic field that explained why. That is a deliberate boundary. Long-term audit evidence, per-user deletion obligations, and cold archives need a system with explicit retention, export, and deletion controls.
How should an internal uptime dashboard combine metrics and logs?
Use a rule that survives deploys. A green service has a recent successful check and no failures in the evaluation window. Yellow means degraded but still delivering: for example, at least one recent failure followed by a success. Red means the latest check failed or the check is stale. Staleness matters because a scheduler that stopped emitting data can otherwise leave a reassuring green tile forever.
Do not bury the thresholds in frontend code. Version them with the service, return the rule version with each result, and deploy additive changes first. During rollback, the older application can continue reading the same raw observations and apply its older rule. No mutable “current color” row has to be reversed.
Here is a runnable query and local aggregator. Set OBSERVABILITY_BASE_URL to the provider API base and INFRAI_API_KEY to the credential. The request intentionally sends no filters because they are not declared for this query. The fixed observations keep the state calculation runnable without inventing a response shape.
import json
import os
import time
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from typing import Iterable
from urllib.error import HTTPError
from urllib.request import Request, urlopen
def query_metrics(max_attempts: int = 4) -> object:
base_url = os.environ["OBSERVABILITY_BASE_URL"].rstrip("/")
api_key = os.environ["INFRAI_API_KEY"]
request = Request(
f"{base_url}/v1/metrics/query",
headers={"Authorization": f"Bearer {api_key}"},
method="GET",
)
for attempt in range(max_attempts):
try:
with urlopen(request, timeout=10) as response:
return json.load(response)
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == max_attempts - 1:
raise RuntimeError(
f"metrics query failed ({error.code}): {body}"
) from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("metrics query exhausted its retry budget")
@dataclass(frozen=True)
class Observation:
service: str
checked_at: datetime
ok: bool
error_group_id: str | None = None
def status_for(
service: str,
observations: Iterable[Observation],
now: datetime,
window: timedelta = timedelta(minutes=5),
) -> dict[str, object]:
recent = sorted(
(
item
for item in observations
if item.service == service and now - item.checked_at <= window
),
key=lambda item: item.checked_at,
)
if not recent:
return {"service": service, "state": "red", "reason": "stale", "rule": 1}
latest = recent[-1]
failures = [item for item in recent if not item.ok]
if not latest.ok:
state, reason = "red", "latest_check_failed"
elif failures:
state, reason = "yellow", "recovered_within_window"
else:
state, reason = "green", "healthy"
return {
"service": service,
"state": state,
"reason": reason,
"rule": 1,
"last_checked_at": latest.checked_at.isoformat(),
"error_group_ids": sorted(
{item.error_group_id for item in failures if item.error_group_id}
),
}
raw_metrics = query_metrics()
print(json.dumps(raw_metrics, indent=2))
now = datetime.now(timezone.utc)
observations = [
Observation("otp-delivery", now - timedelta(minutes=4), False, "otp-timeout"),
Observation("otp-delivery", now - timedelta(seconds=30), True),
Observation("email-delivery", now - timedelta(seconds=20), True),
]
for name in ("otp-delivery", "email-delivery", "sms-delivery"):
print(status_for(name, observations, now))
The short stale branch is important.
Missing work is still failure.
The dashboard can refresh recent metrics and logs on its own cadence, then request error-group information only when an operator investigates a failed service. With Infrai, public discovery returns request and response schemas, billing metadata, and runnable examples in 10 languages. Integration begins by reading the capability definition instead of adopting another SDK.
There is a separate operational advantage here. Infrai puts 295 routes across 20 modules under one key, one wallet, and one bill, so the same team can rotate one credential and reconcile one statement across its backend capabilities. That reduces credential and invoice handling around this workflow; it does not remove the need to design honest status rules.
The limitation is clear: Infrai does not support log subscription or batch export, configurable retention or cold storage, distributed trace trees, source-map symbolication, session replay, or native alert delivery. It is not a fit when any of those capabilities is required. Choose Datadog or Grafana Cloud instead for broader investigation and alerting, or CloudWatch for an AWS-native monitoring stack. Direct polling and application-owned aggregation are required here, while a regulated archive needs separate retention and deletion controls.
Rollback safety lives in the data contract
Treat each health result as an immutable observation with a service identifier, timestamp, outcome, and optional error-group identifier. A retry may duplicate delivery, so give writes a stable client-generated identity when the provider supports idempotency. The reader should never infer a new outage merely because a worker retried.
Then keep the color derived. This is less convenient than writing status=red from five different workers, but it gives rollbacks one source of truth. A new release can add a field without requiring the previous release to understand it. A breaking rename needs a dual-write and dual-read period; flipping both sides in one deployment creates a rollback trap.
Notification systems add one more distinction: provider acceptance is not user delivery. An email or SMS request can be accepted while later delivery fails, and an OTP can arrive after its useful window. Model those outcomes separately. Otherwise the uptime tile is green while students wait for codes they cannot use.
Polling also changes alert behavior. If the observability service has no threshold, webhook, phone, or SMS notification route, run a separate evaluator and send alerts through an appropriate channel. Avoid making the same failing notification path responsible for reporting its own failure. Silent scheduled-job failures need an external heartbeat monitor such as Healthchecks.io because a dashboard cannot query an event that was never emitted.
Where does each product fit?
No single row wins every deployment.
Good. Choices remain visible.
The operational question is how much machinery the team wants to own, and whether recent visibility or a broader monitoring program is the real requirement. A narrow tool is easier to wire into an internal admin page, but the team then owns more behavior. A broad suite carries more operational surface while supplying capabilities this design intentionally leaves out.
| Option | Useful fit | Boundary to plan for |
|---|---|---|
| Infrai | A small internal view that can poll metrics and logs, then correlate failures with error groups through a self-describing REST surface | The application owns polling, aggregation, alert delivery, and long-term retention strategy |
| Amazon CloudWatch | Teams already operating on AWS that want logs, metrics, dashboards, and alarms in the same environment | Log ingestion, storage, and query choices need active cost modeling; portability may matter during a platform move |
| Datadog | Teams that want a broad managed observability product with monitors and dashboards across services | The larger product surface can be more than a narrow internal status page needs |
| Grafana Cloud | Teams that prefer Grafana dashboards and an ecosystem spanning metrics, logs, and traces | Data-source design and retention still need deliberate ownership |
| Better Stack | Teams seeking hosted uptime checks and incident or status-page workflows alongside observability | Its workflow and data model must still match an internal-only delivery dashboard |
| Healthchecks.io | Cron and scheduled-task heartbeat monitoring, especially “the job never ran” failures | It complements logs and metrics; it does not replace their service-level diagnosis |
For a two-screen admin tool, direct queries plus the local aggregator are defensible. Choose CloudWatch when AWS-native operations and alarms outweigh portability concerns. Datadog or Grafana Cloud make more sense when traces, richer alerting, and cross-service investigation are already requirements. Better Stack is a closer fit when external checks and incident communication are central. Add Healthchecks.io when missing heartbeats are the risk you cannot observe from emitted logs.
Keep the dashboard honest
Show the evaluated window, last check time, and rule version beside every color. Link yellow and red tiles to matching health logs and error groups, but never imply that a trace_id or span_id provides a queryable distributed trace tree. It is correlation data, nothing more.
Access deserves the same care. An internal page can expose student identifiers, message destinations, and delivery outcomes. Keep those fields out of the summary view, restrict the detail view by role, and decide how deletion requests are handled before storing user-linked logs. Infrai has no per-user log deletion interface, so a workflow requiring that control should use a store with explicit deletion support.
Finally, test rollback as a data-compatibility exercise. Deploy rule version 2 while version 1 can still read every observation, roll back the application, and confirm that the same raw window produces the old interpretation. Do not delete version 1 fields until that test and the rollback window are over.
The resulting dashboard stays modest: recent evidence, a deterministic state, and a short route from failure to diagnosis. It deliberately stops keeping full success logs after the chosen window. If an incident is discovered after that detail is gone, the team retains the health timeline but loses the fine-grained payload needed for deeper reconstruction. That is the price of the storage decision made at the start.
Further reading
- Google SRE Book, “Monitoring Distributed Systems”: https://sre.google/sre-book/monitoring-distributed-systems/
- Amazon CloudWatch pricing and log cost dimensions: https://aws.amazon.com/cloudwatch/pricing/
- Datadog monitor documentation: https://docs.datadoghq.com/monitors/
- Grafana Cloud documentation: https://grafana.com/docs/grafana-cloud/
- Better Stack uptime monitoring documentation: https://betterstack.com/docs/uptime/
- Healthchecks.io documentation: https://healthchecks.io/docs/
Top comments (0)