DEV Community

UriahHawkins5489
UriahHawkins5489

Posted on

Edtech Observability Stack — Health Monitoring Through Logs, Metrics, and Result Deadlines

An uptime check can prove that an edtech app answers requests while its nightly roster import produces nothing. The useful trade-off is signal quality versus noise: alert on a missed, overdue result backed by a completion record, not on every failed request or quiet interval. Short answer: keep structured logs for investigation, a small set of metrics for trends, explicit job-completion events for freshness, and an external uptime probe for reachability. Page only when the evidence says a promised outcome is missing.

That division works for a small app because each signal has one job. It also avoids pretending that a large telemetry bundle will repair a weak definition of health. A green /health response is infrastructure evidence. It isn't evidence that today's student records arrived.

What should a small-app observability stack monitor for import health?

Start with the user-visible promise. Suppose district files are expected at 02:00 local time, the normal import window is 25 minutes, and downstream results should be available by 03:00. The monitor needs to ask whether a particular scheduled run reached a terminal state and produced a plausible result before that deadline.

This is a state problem, not a log-search problem.

Represent every expected run with a stable key such as tenantId + scheduleDate. Record its scheduled time, start time, completion time, terminal status, input count, accepted count, rejected count, and a correlation ID. Counts aren't proof that the data is correct, but they distinguish “completed with zero source rows” from “worker vanished after start.” That distinction cuts noise immediately: an empty file may be a valid business outcome, while an absent completion event is an operational failure.

Dates are a trap. A district in Los Angeles and a worker running on UTC can disagree about which day owns a 02:00 import, especially around daylight-saving transitions. Create the expectation from the district's configured time zone, persist both the business date and an absolute deadline, and let the evaluator compare absolute timestamps. Do not reconstruct yesterday's expected key from the evaluator's local clock. Otherwise, the monitor can page for a run that was never due while overlooking the run users are actually waiting for; the telemetry stack will faithfully report the wrong model.

Model the promise once.

The resulting data flow is plain. The scheduler creates an expectation. The worker emits structured progress logs and writes one idempotent completion record. A metrics exporter aggregates counters and durations. A separate evaluator compares current time with each expectation's deadline, while an external probe checks that the public app remains reachable. Alerts come from the evaluator and the probe; logs remain supporting evidence.

Build the result monitor first

Here is a compact TypeScript implementation of that core. It deliberately depends on storage and paging interfaces rather than a monitoring vendor. The evaluator also uses an injected clock, which makes boundary tests deterministic.

type ImportStatus = "scheduled" | "running" | "succeeded" | "failed";

type ImportRun = {
  tenantId: string;
  scheduleDate: string;
  dueAtMs: number;
  completedAtMs?: number;
  status: ImportStatus;
  inputRows?: number;
  acceptedRows?: number;
  correlationId: string;
};

type Alert = {
  fingerprint: string;
  summary: string;
  details: Record<string, string | number>;
};

interface RunStore {
  listDueBefore(timestampMs: number): Promise<ImportRun[]>;
}

interface AlertSink {
  open(alert: Alert): Promise<void>;
  resolve(fingerprint: string): Promise<void>;
}

export async function evaluateImports(
  store: RunStore,
  alerts: AlertSink,
  nowMs: number,
): Promise<void> {
  const dueRuns = await store.listDueBefore(nowMs);

  for (const run of dueRuns) {
    const fingerprint = `import-missing:${run.tenantId}:${run.scheduleDate}`;
    const complete = run.status === "succeeded" && run.completedAtMs !== undefined;

    if (complete) {
      await alerts.resolve(fingerprint);
      continue;
    }

    await alerts.open({
      fingerprint,
      summary: "Scheduled roster import has no successful result",
      details: {
        tenantId: run.tenantId,
        scheduleDate: run.scheduleDate,
        dueAtMs: run.dueAtMs,
        status: run.status,
        correlationId: run.correlationId,
      },
    });
  }
}
Enter fullscreen mode Exit fullscreen mode

The fingerprint matters. Running this evaluator every minute should update one incident, not create 60 notifications per hour. Resolution is equally important: once a delayed run succeeds, the same key closes the alert automatically.

One run, one incident.

Keep the write path idempotent too. A worker can retry after losing its response from storage, so uniqueness should be enforced on the expected-run key rather than inferred from how many “finished” log lines appeared. The completion transaction should store the terminal status and counts together. If business validation happens later, give that stage its own expectation and deadline instead of overloading “import succeeded.”

The monitor should expose a few aggregate measurements: expected runs, successful runs, failed runs, overdue runs, and completion duration. OpenTelemetry defines a metric as a measurement captured at runtime and describes instruments including counters and histograms. A counter fits completed-run totals; a histogram fits durations. An overdue value is current state, so an observable gauge is a better semantic match than an ever-increasing counter.

Do not attach raw tenant IDs or correlation IDs as metric attributes. OpenTelemetry's metrics guidance calls out cardinality limits, and identifiers create an unbounded series set. Put those identifiers in the completion record and structured logs. Metrics need bounded dimensions such as import type, environment, and terminal status.

Separate pages from evidence

I would use four signal classes, but only two should normally page a solo operator.

Signal Question answered Retention shape Default action
External probe Can a user reach the app from outside? Small time series Page after confirmed failures
Result freshness Did each promised import finish by its deadline? One record per expected run Page on overdue result
Metrics Is failure rate or duration changing? Aggregated time series Dashboard or threshold alert
Structured logs What happened during this run? Sampled or time-limited events Investigate after an alert

This separation prevents a common cascade. A database slowdown may generate request errors, retries, worker exceptions, latency spikes, and a failed uptime probe. Paging on every symptom multiplies noise without adding decisions. Route them through stable fingerprints, group related symptoms, and let the user-facing breach carry the highest severity.

Silence needs careful handling. “No successful result by 03:00” is actionable because the scheduler created a dated expectation. “No logs for an hour” is ambiguous: perhaps no import was scheduled, perhaps log delivery failed, or perhaps the worker never launched. The expected-run ledger turns silence into a testable condition.

Quiet is not healthy.

There is one more clock to watch. If the evaluator shares the worker's host, a machine outage can stop both the work and the alarm. Run the evaluator in a separate failure domain where practical, and use an external probe for the public endpoint. Configure confirmation across multiple attempts so a short network interruption does not wake anyone, but keep the confirmation window inside the business deadline. The exact attempt count belongs to the service objective; there is no honest universal number.

Choose components by boundaries, not brand count

A small app still needs storage, collection, visualization, and notification, but those roles do not have to arrive as one suite. A useful selection exercise starts with the interfaces: structured JSON logs to standard output, OpenTelemetry metrics, durable completion records in the application database, and a generic alert sink. That makes replacement possible without rewriting the import worker.

Hosted suites and focused tools draw those boundaries differently. Sentry centers its documentation around errors and performance, while Datadog documents separate ingestion and indexing dimensions for logs. Grafana's documentation presents Loki for logs and Prometheus-compatible metrics integrations. These are objective scope and data-model differences, not a ranking. If evaluating products, compare at least those three shapes against the same workload and retention policy rather than counting dashboard features.

The cost question is mostly a volume-design question. Error stack traces, retry loops, and per-row logs can expand far faster than five bounded metrics and one completion record per run. Avoid logging every student row. Log a summary plus the correlation ID, retain rejected-row details in access-controlled application storage when the product genuinely requires them, and treat student data as sensitive. Cost control and privacy point in the same direction here.

Run a short replay test before committing to any backend. Send representative successes, a retry storm, an empty-but-valid file, a malformed file, and a run that starts but never completes. Then verify ingestion delay, query ergonomics, alert deduplication, retention controls, exportability, and the behavior when the telemetry destination is unavailable. Telemetry failure must not convert a valid import into a failed import; buffer within a strict bound and drop low-value diagnostics before blocking the job.

Test the absences that matter

Happy-path tests are insufficient because this monitor exists to detect missing state. Use a fake clock and cover the boundary immediately before and at dueAtMs. Test a never-started run, a running job that crosses its deadline, an explicit failure, a late success that resolves an incident, and repeated evaluations that preserve one fingerprint. Also test two tenants due at the same time so deduplication cannot collapse distinct failures.

Deployment deserves a synthetic import that contains no real student data. Schedule it through the same path as production work, give it a known small result, and verify the completion record. This checks more than a generic HTTP probe without leaking personal information. Keep it separate from customer freshness alerts because its failure means “the pipeline may be broken,” not “a specific district result is late.”

Operationally, define ownership before turning paging on. The runbook should say how to find a run by correlation ID, distinguish a valid zero-row input from a missing source, retry idempotently, and confirm downstream publication. Review overdue alerts after schedule changes and school-calendar exceptions; stale expectations create believable but useless pages. Finally, inspect telemetry volume and dropped-event counts during retry tests, then verify that an outage in the observability path cannot halt imports.

The right stack is the smallest collection of replaceable components that preserves those behaviors. Health is an outcome with a deadline, not a green process and not a pile of searchable logs. Once that model is correct, backend selection becomes a constrained implementation choice: bounded metrics for trends, structured events for diagnosis, durable expectations for silence, and an independent probe for reachability.

Sources

Top comments (0)