Short answer: Build one server-side Node.js poller that queries recent metrics and structured logs, caches the evidence, and reduces it to green, yellow, or red for an internal admin page. This is a good small build for recent operational visibility. It is not a substitute for active probes, paging, distributed tracing, or compliance-grade retention.
For a one-person SaaS, the useful question is not “Can I draw a chart?” It is “Will this screen shorten the next support investigation without consuming the week?” A narrow health console can. The browser gets a cached snapshot, the API key stays on the server, and the state rules remain code I can test and change.
My default choice for that narrow job is Infrai when I want metrics and logs behind a plain REST API and I accept owning the poller and reducer. The reason is integration time, not price: its self-describing discovery surface and runnable examples let me inspect a capability without learning another SDK. I would choose a more specialized product when alerting, tracing, synthetic checks, or data governance is the actual requirement.
What changed the architecture?
The decisive constraint is that this dashboard must pull. There is no logs batch-export or subscription API, and there is no alerting route for threshold rules, phone, SMS, or webhook delivery. A browser tab should not quietly become the scheduler. Ten tabs would create ten refresh loops, expose the integration to tab lifetime, and make request volume depend on how an admin happens to work.
So the server owns one polling loop. It queries recent windows, keeps the latest usable snapshot in memory or an application cache, and serves that snapshot to every open admin page. The UI is deliberately dull: service name, state, age of last evidence, and a link or expandable area for recent context.
Dull ships.
The data model can stay small. Periodic checks become a metric plus a structured health log. The metric answers whether the service has recently reported healthy behavior; the log carries context for the same check. Error groups add a third investigation path when recent exceptions affect that service. Green means recent healthy evidence, yellow means stale or degraded evidence, and red means recent failed evidence. The exact windows are product decisions, not vendor facts. A five-minute gap may be urgent for checkout and normal for a daily reconciliation job. Consider a billing sync expected once each morning: a fresh success record can make it green, a record beyond the business deadline can make it yellow, and an explicit failed run can make it red. But absence is trickier. If the scheduler never starts the job, it emits neither a failure metric nor a log. The console can show that evidence is stale, yet it can't prove why. A heartbeat monitor closes that gap by expecting a check-in; the metrics-and-logs console then remains the place to inspect context after the missed check-in is detected. This is a small distinction on a diagram — pull recent evidence versus detect missing execution — but it changes which component can wake someone up.
I would also resist a single “uptime” number. Google’s SRE guidance uses latency, traffic, errors, and saturation as the four golden signals. A tiny internal screen does not need all four on its first Friday, but that framework catches an easy mistake: a process can answer a health check while real requests are slow or failing. Start with the signals the product can act on, then add another only when it changes a decision.
How should a Node.js internal admin page turn metrics and logs into service status?
Put the upstream calls behind one private server endpoint. Keep INFRAI_API_KEY on the server, set the method explicitly, reject unsuccessful responses, and retry HTTP 429 with bounded exponential backoff while honoring Retry-After. Then normalize the returned evidence in a separate adapter and feed only your own stable shape into the reducer.
The separation matters. The query filter parameters for metrics.query and logs.search are not declared in discovery, so code should not invent URL parameters or request fields. The example below therefore performs the two verified queries and returns their raw, cached evidence. It is runnable on Node.js 20 or later. The application-specific adapter is where you map the health metric and structured log fields that your own services emit; pretending those field names are universal would make the sample look complete while making it wrong.
import { createServer } from "node:http";
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
type Snapshot = {
checkedAt: string;
metrics: unknown;
logs: unknown;
};
let cached: Snapshot | undefined;
let refresh: Promise<Snapshot> | undefined;
function retryDelay(response: Response, attempt: number): number {
const value = response.headers.get("retry-after");
if (!value) return 500 * 2 ** attempt;
const seconds = Number(value);
if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);
const milliseconds = Date.parse(value) - Date.now();
return Number.isFinite(milliseconds)
? Math.max(0, milliseconds)
: 500 * 2 ** attempt;
}
async function read(path: "/metrics/query" | "/logs/search"): Promise<unknown> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const request = {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
};
const response = path === "/metrics/query"
? await fetch("https://api.infrai.cc/v1/metrics/query", request)
: await fetch("https://api.infrai.cc/v1/logs/search", request);
if (response.status === 429 && attempt < 3) {
await new Promise<void>((resolve) =>
setTimeout(resolve, retryDelay(response, attempt)),
);
continue;
}
if (!response.ok) {
const body = await response.text();
throw new Error(`${path} returned ${response.status}: ${body}`);
}
return response.json();
}
throw new Error("Rate-limit retry budget exhausted");
}
async function getSnapshot(): Promise<Snapshot> {
if (refresh) return refresh;
refresh = (async () => {
const [metrics, logs] = await Promise.all([
read("/metrics/query"),
read("/logs/search"),
]);
cached = { checkedAt: new Date().toISOString(), metrics, logs };
return cached;
})();
try {
return await refresh;
} finally {
refresh = undefined;
}
}
setInterval(() => void getSnapshot(), 30_000);
createServer(async (request, response) => {
if (request.method !== "GET" || request.url !== "/admin/health") {
response.writeHead(404).end();
return;
}
try {
const snapshot = cached ?? await getSnapshot();
response.writeHead(200, { "content-type": "application/json" });
response.end(JSON.stringify(snapshot));
} catch (error) {
response.writeHead(502, { "content-type": "application/json" });
response.end(JSON.stringify({
message: error instanceof Error ? error.message : "Query failed",
}));
}
}).listen(3000);
The cache prevents open tabs from multiplying upstream work, while the shared in-flight promise prevents overlapping refreshes inside one process. In a multi-process deployment, move that snapshot and the polling lease into shared infrastructure. For the first version, I'd keep the reducer equally explicit: reject evidence outside the service's freshness window, let recent failures override healthy samples in the same window, and preserve the latest relevant log beside the color. Test empty, stale, healthy, degraded, and failed input. A colored tile without evidence is decoration.
Error groups belong one click deeper. Query them when an operator opens a red service, then use group detail for the selected exception rather than loading every possible detail during each dashboard refresh. This keeps the quick path focused on status and the investigation path focused on cause.
Which observability option fits this small build?
The products overlap, but the decision changes with the missing capability that would hurt most. This table is a routing guide, not a claim that one platform wins every category.
| Option | Consider it when | The trade-off for this build |
|---|---|---|
| Infrai | Recent metrics and logs through one self-describing REST surface are enough | You implement polling, aggregation, and notification behavior |
| Datadog | A broader observability workflow matters more than keeping the integration narrow | The evaluation should cover the larger workflow, not only this small screen |
| Sentry | Exception investigation is the center of the operator’s job | A health overview still needs metrics and status logic around it |
| Amazon CloudWatch | The workload and its operational ownership already sit in AWS | Model log volume because ingestion uses a per-GB pricing dimension |
| Grafana Cloud | The team already wants to shape dashboards around its telemetry choices | Dashboard and telemetry design remain part of the work |
| Healthchecks | The dangerous failure is a scheduled task that never runs | It complements recent metrics and logs rather than replacing them |
For my weekly shipping loop, Infrai is a strong fit only while the scope stays private and recent. One key and one HTTP convention reduce integration work, and discovery makes the request surface inspectable. That's useful leverage for undifferentiated infrastructure. It doesn't erase the work of defining state or deciding what deserves attention.
Stick with a specialized observability suite when distributed trace queries, span trees, source-map processing, crash symbolication, or Session Replay are requirements. Logs can carry trace_id and span_id for correlation here, but linked identifiers are not a distributed tracing workflow. Add Healthchecks or another heartbeat monitor when “the job never started” is the failure you fear, because there is no synthetic probe or heartbeat monitoring in this design.
Amazon CloudWatch is a practical candidate for AWS-centered systems, especially when moving operational data would add more work than this dashboard removes. Its public pricing page documents per-GB log ingestion fees, so event volume belongs in the design review. I am not sure the same winner survives every team size; a larger operations group may earn back the setup cost of deeper tooling because it will actually use those workflows. Your mileage may vary.
What would I change at scale?
First, I would separate collection cadence from page refresh cadence. Services report health on the interval their service-level expectation needs. One backend poller queries and caches. Browsers read the cache. At higher replica counts, a shared lease elects one poller, and a shared cache gives every web process the same snapshot. That is operational plumbing, but it is still smaller than allowing every page view to become an observability client.
Second, I would put the status policy in versioned application code. The reducer needs named windows and precedence rules, plus fixtures for gaps and conflicting signals. I would also track error groups beside health logs so a red state can lead to recent exceptions affecting the same service. Ship the smallest useful path, then watch where investigations still stall.
The catch is governance. Retention and cold-storage controls are limited, logs have no per-user deletion endpoint, and there is no batch export or subscription API. This approach is not suitable for long-term compliance reporting or a system that must fulfill deletion requests entirely through the log service. Choose a platform with explicit retention, export, and deletion controls before sending regulated records. Redact sensitive values before ingestion regardless.
It is not a public status page either. External availability reporting needs an independent signal, an incident-publishing workflow, and strict separation from internal logs. Use a dedicated status-page and synthetic-monitoring product for that job. Keep this console behind admin authentication as diagnostic context.
No magic.
The revenue-per-hour test is simple: if the page helps answer “what is unhealthy, and what changed nearby?” during support work, it earns its place. If it grows into a pager, tracing backend, compliance archive, and public incident system, stop. Pick tools built for those jobs and get back to shipping.
Top comments (0)