TL;DR: use an external uptime monitor to answer whether the notification service is reachable, and use an internal health dashboard built from structured delivery events to reconstruct why a player never received a message. A green health endpoint cannot establish that a queued guild invite reached its destination. For a beginner running a Node.js Express service, the practical choice is both signals with sharply different jobs, not one elaborate screen pretending to cover every failure.
The deciding constraint is incident reconstruction. When a tournament reminder disappears, I want a short chain of evidence: accepted, queued, attempted, acknowledged or failed. The public probe is one entry in that chain. It is deliberately shallow because coupling it to every downstream dependency turns a partial delivery problem into a declared outage.
Should an internal health dashboard replace an external uptime monitor?
A process can answer an HTTP request while its worker is stalled, its queue is growing, or a destination is rejecting sends. The check proves reachability at a particular path from a particular vantage point. It does not prove end-to-end notification delivery.
That limitation matters.
Keep the route boring. Return success only after startup has completed and the process is ready to accept work. Put diagnostic detail in telemetry rather than the response body; dependency names and failure detail are useful to an operator, but they need not be exposed to every caller.
import express from "express";
const app = express();
let ready = false;
app.get("/health", (_request, response) => {
response.status(ready ? 200 : 503).json({ status: ready ? "ok" : "starting" });
});
async function start(): Promise<void> {
// Complete required startup work before accepting notification traffic.
ready = true;
app.listen(3000);
}
void start();
This endpoint supports a binary external observation. It should not query the notification provider, drain a queue, or send a test message on every poll. Those actions have different cost, latency, and side-effect profiles. Mixing them makes the result harder to interpret.
A useful probe records timestamp, target, region or vantage point, duration, status code, and timeout category. That is enough to answer a narrow question: could an independent caller reach the service then?
This split is not suitable for every stage. A single-process prototype with no asynchronous delivery has little incident trail to reconstruct, so a basic uptime check and structured application logging may be enough. At the other extreme, a service with several queues, workers, and regions may need distributed traces and automated correlation rather than a hand-built dashboard. The trade-off is extra event storage and correlation work in exchange for evidence that survives process boundaries.
The failed shortcut: one dashboard, one status
The simple design maps process health to delivery health. It is attractive because the dashboard has one large indicator and an incident begins with a clean yes-or-no answer. It fails as soon as work becomes asynchronous. HTTP acceptance, queue processing, destination response, retry scheduling, and final outcome happen at different times and may cross different processes.
The reverse shortcut also fails. An external monitor can detect total unreachability from outside the deployment, but it cannot infer that only clan invitations are stuck while password alerts continue. Internal logs can show that distinction, yet logs produced by the same unavailable process cannot independently prove public reachability.
That gives the two surfaces separate ownership:
| Signal | Question it answers | Keep out of it |
|---|---|---|
| External probe | Can a caller reach the service now? | Queue depth and recipient-level detail |
| Delivery events | What happened to this notification? | Global claims based on one event |
| Aggregated metrics | Is failure or latency changing by channel and message class? | High-cardinality recipient identifiers |
No signal wins. The join does.
Build a reconstruction trail, not a log pile
For each notification, create an opaque operation ID at acceptance and carry it through queue metadata and delivery events. Record a stable message class such as tournament_reminder, the channel, attempt number, outcome, duration, and a normalized failure category. Avoid putting message bodies, access tokens, or raw recipient addresses into telemetry. The operation ID lets an authorized operator follow one delivery without making personal data the primary index.
type DeliveryOutcome =
| "accepted"
| "attempted"
| "delivered"
| "retry_scheduled"
| "failed";
type DeliveryEvent = {
operationId: string;
messageClass: "tournament_reminder" | "guild_invite" | "security_alert";
channel: "push" | "email";
attempt: number;
outcome: DeliveryOutcome;
durationMs?: number;
failureCategory?: "timeout" | "rate_limited" | "rejected" | "invalid_destination";
occurredAt: string;
};
function recordDelivery(event: DeliveryEvent): void {
process.stdout.write(`${JSON.stringify(event)}\n`);
}
A focused experiment can use three synthetic operation IDs and four controlled outcomes: accepted, delivered, retry scheduled, and terminal failure. The goal is not to manufacture a benchmark. It is to verify that an operator can reconstruct ordering across the web process and worker, distinguish retryable from terminal outcomes, and see a missing transition. If accepted exists but no later event arrives within the service's own delivery objective, the dashboard should expose incomplete work rather than repainting the health route red.
Metrics should stay aggregate. Count attempts and outcomes by message class, channel, and failure category; measure queue age and delivery duration with bounded dimensions. Keep operation IDs in logs or traces, where lookup is intentional. Putting every operation ID into a metric label creates a dimension per notification and makes the aggregate surface expensive without improving the incident decision.
Sampling needs care here. OpenTelemetry distinguishes head sampling, decided before the trace completes, from tail sampling, decided after all or most spans are available. A low head-sampling rate can discard the rare failed delivery that matters most. Tail sampling can retain traces based on completed outcomes, but it requires a collector to assemble spans and delays the decision. For a solo-operated service, I would first preserve compact terminal delivery events, then add traces only where cross-process timing remains ambiguous. Ship the evidence you can afford to keep.
Start small.
Make changes reversible during the incident
Observability does not reduce impact by itself. The delivery path needs a control that can disable a risky message class, switch a new worker behavior off, or reduce traffic without rebuilding the application. Feature toggles provide that separation between deployment and release, as Martin Fowler describes, but they also introduce configuration that must be managed and retired.
Record toggle state or delivery-strategy version on each event. Then an operator can compare failures before and after a change instead of guessing which code path handled an attempt. Do not log the entire configuration object; a stable version or small set of relevant states is easier to query and less likely to leak unrelated settings.
The rollback rule should be written before rollout. For example: if the new worker path produces terminal failures in the synthetic test, disable that path and verify that new operations carry the previous strategy version. The numbers here define a test fixture, not a universal reliability target. Production thresholds must come from expected traffic, acceptable player impact, and enough baseline data to separate an incident from ordinary variation.
What should you measure before copying this design?
Measure the questions your incident process actually asks. Start with time to locate an operation, percentage of accepted notifications that reach a terminal event, age of the oldest nonterminal operation, delivery duration by message class, and the external probe's success and duration from each configured vantage point. Also measure telemetry volume per delivered notification. Token and storage costs are operating constraints, even when the first dashboard is tiny.
Run a game-day test with a delayed worker, a rejected destination, and an unreachable web process. The expected evidence should differ in all three cases. If the same red indicator represents each one, the model is still too coarse. If reconstruction requires searching free-form messages by player address, the event contract is too weak.
Choose retention and sampling after observing event volume and incident lookup needs. Choose alert thresholds after collecting a baseline. Choose an external check interval from the detection objective and acceptable probe traffic, not from a dashboard default. These decisions are local; copying somebody else's intervals gives precise-looking numbers with no connection to player impact.
The final decision is modest: keep the independent reachability check small, make delivery events complete enough to reconstruct one operation, and aggregate only dimensions that drive an action. A dashboard is then a view over evidence rather than the source of truth. That distinction is what lets a small team debug a missed tournament reminder without turning every downstream wobble into a full-service outage.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.