TL;DR: For a small Node.js media SaaS, the best observability stack is a quiet pipeline, not the longest shopping list. Probe the public health endpoint from Europe and the US, emit one structured event for each AI agent loop, derive latency and cost metrics from those events, and send exceptions to a separate error stream. Page only on sustained user-visible failure. Everything else belongs in a dashboard or a daily review.
That split gives each signal one job. A probe answers whether readers can reach the service. A loop event explains where generation time and usage went. An error record preserves debugging context. None of those facts, by itself, proves that a person needs to be interrupted.
For a one-person SaaS, interruption is part of infrastructure cost. I use a revenue-per-hour lens: a signal earns its keep when it changes the next engineering decision. Shipping weekly matters more than maintaining five overlapping collectors. Outsource undifferentiated storage and notification plumbing when that saves operator time, but keep the event schema and alert rules portable.
What should a small observability stack monitor at the health endpoint?
Start with the user-visible contract. An external probe should call a cheap health route through the same public path as normal traffic. Run it from at least one location in Europe and one in the US because the application serves both regions. Keep the response narrow: process readiness and dependencies required to serve the next request. Do not run a model call from the health handler. That would mix a cheap availability check with a slow, variable dependency and could create spend during an incident.
The probe records status, elapsed time, location, and a check identifier. Alerting comes later. One failed sample is evidence, not a page. A practical policy requires repeated failures or a rolling-window breach, then confirms that the failure is user-visible. Derive the exact threshold from the service objective and probe cadence rather than copying somebody else's number.
Silence has value.
Google's SRE guidance separates symptoms from causes and describes latency, traffic, errors, and saturation as the four golden signals. That distinction matters here. An outside-in failed check is a symptom. A database pool at capacity is a possible cause. Page on the symptom when it is sustained; retain cause signals for diagnosis.
The AI loop needs different treatment. For every run, capture total duration, duration by step, provider-reported model usage, outcome, and a stable correlation ID. Compute cost after ingestion using the applicable model contract. Keeping raw usage separate from the rate table avoids baking changeable prices into application events.
Consider a concrete review path. The European probe is healthy, the US probe is healthy, and completed media jobs are still arriving, but the loop-latency distribution has shifted. That is not paging evidence. Open the slow runs by correlation ID, compare step durations, check whether retries explain extra usage, and decide during working hours whether the change harms the publishing promise. If both public probes instead show sustained failures, loop detail becomes diagnostic context after the page. This ordering keeps a performance investigation from impersonating an availability incident, while still preserving enough detail to find the slow stage.
| Signal | Keep | Review | Interrupt |
|---|---|---|---|
| External probe | Status and latency by region | Availability trend | Sustained user-visible failure |
| Agent loop | Step timing, usage, outcome | Tail latency and cost per successful output | Only when tied to a service objective |
| Application log | Structured operational event | Investigation or sampled review | Never from arbitrary message text |
| Exception | Type, stack, release, correlation ID | New and recurring groups | Only for material user impact |
The table is intentionally small. More telemetry is easy. Useful attention is hard.
Keep it boring.
The constraint that changed the design
A media agent loop is not one request in the useful sense. It may fetch source material, call a model several times, validate an answer, and persist an article. A single HTTP duration hides the slow step. Logging every prompt and response, however, raises volume and can retain editorial or user data that the operator never needed for latency analysis.
So the unit of observation should be the loop and its named steps, not every internal function. Emit dimensions with bounded values: operation, deployment environment, region, model identifier, and outcome. Put unique values such as runId in logs or trace context, not metric labels. Otherwise every run creates a new time series and turns a useful metric into high-cardinality noise.
Health checks and loop measurements answer different questions. A model call can become slower while the API remains available. That is a product-performance review item until it violates a declared user-facing objective. Conversely, a fast background loop says nothing about whether DNS, TLS, routing, or the public server path works from Europe.
Use three stores conceptually, though one backend may fill more than one role: a metrics store for aggregate distributions, a structured event store for per-run investigation, and an indexed error stream for stack traces and grouping. Test retention controls, export format, regional ingestion, query latency at the expected volume, and the work required to restore dashboards after migration. An analytical database can serve event queries; a managed log service or another columnar store can fill the same role. The requirement is queryable structured events, not a particular logo.
The smallest working Node.js implementation
The application can own a compact event contract and send it through a generic sink. This TypeScript uses standard Node.js timing and emits JSON. It assumes no collector or storage vendor.
import { performance } from "node:perf_hooks";
import { randomUUID } from "node:crypto";
type LoopEvent = {
kind: "agent_loop";
runId: string;
region: "eu" | "us";
model: string;
outcome: "ok" | "error";
durationMs: number;
inputTokens?: number;
outputTokens?: number;
steps: Record<string, number>;
occurredAt: string;
};
type EventSink = (event: LoopEvent) => Promise<void>;
export async function measureLoop<T>(
region: LoopEvent["region"],
model: string,
sink: EventSink,
run: (record: (name: string, ms: number) => void) => Promise<{
value: T;
usage?: { inputTokens: number; outputTokens: number };
}>,
): Promise<T> {
const started = performance.now();
const runId = randomUUID();
const steps: Record<string, number> = {};
let outcome: LoopEvent["outcome"] = "ok";
let usage: { inputTokens: number; outputTokens: number } | undefined;
try {
const result = await run((name, ms) => { steps[name] = ms; });
usage = result.usage;
return result.value;
} catch (error) {
outcome = "error";
throw error;
} finally {
await sink({
kind: "agent_loop", runId, region, model, outcome, steps,
durationMs: Math.round(performance.now() - started),
inputTokens: usage?.inputTokens,
outputTokens: usage?.outputTokens,
occurredAt: new Date().toISOString(),
});
}
}
Keep the sink off the critical path in production by buffering events and applying a bounded flush timeout. Decide what happens when telemetry delivery fails: the media job should normally continue, while a local counter records dropped events. Observability must not become a new availability dependency.
Cost calculation belongs downstream. Join token usage to a versioned rate table, preserve currency and effective date, and calculate cost per successful output as well as per attempted loop. Failed retries then remain visible. A rate change can apply prospectively without rewriting the original event.
Test the pipeline before trusting it
Start with the event contract. Assert that a successful loop records duration, outcome, and usage. Assert that a thrown exception still produces an error outcome. Assert that sink failure follows the chosen delivery policy. A fake sink keeps the unit suite off the network.
Then test alert behavior with synthetic time series: one regional failure, several consecutive failures, slow-but-successful probes, and missing samples. Define the expected state for each case. Missing telemetry is not automatically a failed service, but a silent probe fleet deserves its own non-paging diagnostic.
Privacy is part of the design. Do not put article bodies, prompts, access tokens, or raw model responses in the default loop event. If debugging requires captured content, make it a separate, access-controlled workflow with an explicit retention period. Metadata is the useful default.
Content is not a metric.
Review a weekly scorecard: successful outputs, failed outputs, total and step latency distributions, token usage, computed cost, dropped telemetry, and probe results by region. That cadence fits weekly shipping. It catches gradual drift without converting every fluctuation into an interruption.
What I would change at scale
At higher volume, move from one event per completed loop toward spans for step relationships, histogram metrics for fleet-wide latency, and sampled structured events for examples. OpenTelemetry defines vendor-neutral telemetry APIs and conventions. Adoption still has a cost: context propagation, collector operations, schema governance, and sampling rules need ownership. Add it when cross-service correlation pays for that work.
Separate analytical and operational paths. Paging needs short delays and few dependencies. Analysis can batch, compress, enrich, and retain data for longer queries. A single backend may receive both, but the application should not assume identical delivery guarantees.
Scale changes cardinality decisions before it changes dashboard design. Enforce an allowlist of metric dimensions, version the event schema, and reject accidental payload growth in continuous integration. Longer retention is not automatically better. It increases storage, privacy, and governance work.
The trade-off remains signal quality versus noise. Spend attention on events that alter a shipping decision or reveal user harm. Everything else can wait for weekly review.
Ship the useful signal.
Top comments (0)