DEV Community

BriarVoss47291
BriarVoss47291

Posted on

Cheap Python Uptime Dashboard: Custom Metrics for Internal Admin Status

Short answer: a cheap Python uptime dashboard should use custom metrics to detect missing marketplace results, not treat a green process as proof that scheduled work happened. Record one durable row for each run and render the same facts on the internal admin status page. During a rollback, keep the writer schema backward-compatible so old and new workers can report liveness at the same time. That choice matters more than dashboard polish.

A green /health response proves that a process can answer. It does not prove that the 02:00 supplier feed was fetched, parsed, accepted, and published. For this job, the decisive signal is the age of the last successful run that produced an expected result, qualified by whether a run is currently active.

What custom metrics should a cheap uptime dashboard show?

The awkward case is an import that completes with no listings. Zero can mean a quiet supplier, a broken upstream file, an accidental filter, or a parser that accepted nothing. A single success counter erases those distinctions.

Zero is ambiguous.

Define the contract before writing the dashboard. A run has a scheduled timestamp, start and finish timestamps, outcome, input count, accepted count, rejected count, and a reason code. Keep reason codes bounded: completed, empty_expected, empty_unexpected, fetch_failed, parse_failed, and publish_failed are useful dimensions; raw exception messages are not. Prometheus explicitly warns against high-cardinality labels, so supplier IDs, filenames, trace IDs, and exception text belong in logs or the run record rather than metric labels.

The decision rule can stay small. Alert when no qualifying run has finished within the schedule interval plus a grace period. Treat empty_expected as qualifying only when an explicit business rule predicted an empty feed. Suppress the missing-result alert while a run is active and still inside its maximum duration, but raise a separate stuck-run alert after that bound. This separates lateness from failure.

That split matters.

For example, a daily feed expected by 02:00 might use a 45-minute arrival grace period and a 90-minute run limit. Those are example policy values, not universal defaults. Derive them from the supplier agreement and observed duration distribution, then put them under review like any other production threshold.

Build the run ledger before the chart

The data flow is plain: the scheduler creates a run record, the worker updates it at phase boundaries, a metrics collector translates completed records into low-cardinality time series, and an internal page reads a summarized endpoint. Alerts query the metrics backend. The ledger remains the audit trail when a chart has been aggregated or a process restarts.

Here is a runnable in-memory version of that core. It models the contract and the alert decision without coupling the worker to a dashboard library. A production adapter would replace RunStore with a transactional database implementation and emit the returned measurements through the metrics client already used by the service.

from __future__ import annotations

from dataclasses import dataclass, replace
from datetime import datetime, timedelta, timezone
from enum import Enum
from threading import Lock
from typing import Iterable
from uuid import uuid4


class Outcome(str, Enum):
    RUNNING = "running"
    COMPLETED = "completed"
    EMPTY_EXPECTED = "empty_expected"
    EMPTY_UNEXPECTED = "empty_unexpected"
    FAILED = "failed"


@dataclass(frozen=True)
class ImportRun:
    run_id: str
    feed: str
    scheduled_at: datetime
    started_at: datetime
    finished_at: datetime | None = None
    outcome: Outcome = Outcome.RUNNING
    input_count: int = 0
    accepted_count: int = 0
    rejected_count: int = 0


class RunStore:
    def __init__(self) -> None:
        self._runs: dict[str, ImportRun] = {}
        self._lock = Lock()

    def start(self, feed: str, scheduled_at: datetime) -> ImportRun:
        run = ImportRun(
            run_id=str(uuid4()),
            feed=feed,
            scheduled_at=scheduled_at,
            started_at=datetime.now(timezone.utc),
        )
        with self._lock:
            self._runs[run.run_id] = run
        return run

    def finish(
        self, run_id: str, outcome: Outcome, input_count: int,
        accepted_count: int, rejected_count: int,
    ) -> ImportRun:
        if outcome is Outcome.RUNNING:
            raise ValueError("a finished run cannot remain running")
        if min(input_count, accepted_count, rejected_count) < 0:
            raise ValueError("counts cannot be negative")
        with self._lock:
            current = self._runs[run_id]
            finished = replace(
                current,
                finished_at=datetime.now(timezone.utc),
                outcome=outcome,
                input_count=input_count,
                accepted_count=accepted_count,
                rejected_count=rejected_count,
            )
            self._runs[run_id] = finished
            return finished

    def all(self) -> Iterable[ImportRun]:
        with self._lock:
            return tuple(self._runs.values())


def evaluate_liveness(
    runs: Iterable[ImportRun], now: datetime, freshness: timedelta,
    max_runtime: timedelta,
) -> dict[str, float | int]:
    run_list = list(runs)
    qualifying = [
        run for run in run_list
        if run.finished_at is not None
        and run.outcome in {Outcome.COMPLETED, Outcome.EMPTY_EXPECTED}
    ]
    active = [run for run in run_list if run.outcome is Outcome.RUNNING]
    last_finish = max(
        (run.finished_at for run in qualifying if run.finished_at is not None),
        default=None,
    )
    age = (now - last_finish).total_seconds() if last_finish else float("inf")
    stuck = sum(now - run.started_at > max_runtime for run in active)
    return {
        "last_qualifying_result_age_seconds": age,
        "active_runs": len(active),
        "stuck_runs": stuck,
        "missing_result": int(age > freshness.total_seconds() and not active),
    }


if __name__ == "__main__":
    now = datetime.now(timezone.utc)
    store = RunStore()
    run = store.start("catalog_daily", now - timedelta(minutes=20))
    store.finish(run.run_id, Outcome.COMPLETED, 814, 803, 11)
    print(evaluate_liveness(
        store.all(), now, timedelta(hours=25), timedelta(minutes=90)
    ))
Enter fullscreen mode Exit fullscreen mode

The output includes four signals, but only three need to become gauges. missing_result is derived state and can be an alert expression instead of another stored metric. That keeps policy in one place. The internal page can show the latest ten ledger rows, the current age, and the threshold; it does not need to become a second monitoring system.

One trap deserves emphasis: do not update a last_success timestamp before the publish transaction commits. A worker that parses 803 valid listings and then fails to publish them produced no marketplace result. Advance the timestamp after the externally meaningful boundary, using a database transaction or an idempotent finalization step.

This approach has limits. It is unsuitable when an external observer must prove public reachability, DNS resolution, certificate validity, or a user's full browser journey; use an independent probe for those checks. A ledger also adds a database write to every run and requires retention rules. The trade-off is deliberate: an external probe sees availability from outside but cannot establish that a scheduled catalog publication produced valid rows, while the ledger sees business completion but can share a failure domain with the importer. For a high-risk feed, use both signals and keep their alerts independent.

Make rollback a data-contract exercise

Instrumentation changes often ship beside worker changes, so rollback safety depends on mixed-version behavior. The run record should be append-friendly: add nullable fields, tolerate unknown reason codes in readers, and avoid renaming an existing outcome in the same deployment that changes worker logic. Deploy the reader first, then the writer, and remove old fields only after the rollback window closes.

Keep metric names stable during that window too. If a new worker replaces import_last_success_timestamp_seconds with a differently defined series, emit both temporarily and make the alert accept either. The compatibility period costs a little duplication. It buys a rollback that does not turn monitoring blind exactly when confidence is lowest.

Duplication is temporary.

The alert must survive the application rollback. Run alert evaluation outside the import worker, and calculate liveness from durable facts. A restarted process loses in-memory counters; the ledger lets the collector reconstruct gauges. Prometheus counters can reset on restart by design, while gauges represent values that may move up or down, so a Unix timestamp or current active-run count fits a gauge.

The same discipline helps an eval-driven AI pipeline. Treat parser or classifier revisions like prompt revisions: replay fixed feed fixtures, assert outcome and count invariants, and compare the old and new decision rules before changing an alert. Token cost may matter if an import enriches listings with a model, but it is a separate metric from liveness. Record aggregate model calls and tokens for a run; never let a cost dashboard decide whether the catalog was published.

Test the silence, not only the success path

A normal unit test proves that finish stores counts. The valuable tests advance a fake clock across boundaries: no prior run, late but active, active beyond maximum runtime, expected zero, unexpected zero, failure after parsing, and two overlapping runs. Property tests can enforce that a failed run never makes result age younger and that completing a qualifying run never makes it older.

Then rehearse a rollback with both schemas. Start a run using the new writer, read it with the old-compatible collector, roll the worker back, finish another run, and verify that the alert remains evaluable. This is closer to an eval harness than a screenshot test: fixed input, explicit expected decision, repeatable result.

Do not page on every rejected listing. Use a ratio or count warning with a minimum-volume guard, and keep it below the urgency of a missing publication. Otherwise a feed containing one rejected row out of one input looks catastrophic while a completely absent feed waits quietly.

Silence is worse.

Operate the smallest useful surface

The page should answer three questions without interpretation: when did the last qualifying result finish, is a run active or stuck, and what happened in recent runs? Show timestamps in UTC with an explicit label, plus local display only as a convenience. Link each row to correlated logs by run ID, but do not put run IDs into metric labels.

Authentication and authorization still apply because an internal panel exposes supplier names, schedules, and failure context. Keep mutation controls out of the first version. A read-only page has a smaller failure and privilege surface, while retries and manual imports can remain in an audited operations path. Cache the summary briefly and set a query timeout so a slow metrics store cannot consume the admin application's worker pool.

The operational checklist is short enough to write as prose. Confirm that every scheduled run creates a ledger entry, every terminal path finalizes it, and the publish boundary defines success. Verify alerts with a fake clock, a deliberately stuck run, an unexpected empty feed, and a stopped scheduler. During deployment, watch both metric schemas through the rollback window. Finally, make the on-call runbook name the feed owner, the expected arrival window, the safe retry mechanism, and the evidence required before suppressing an alert.

This design stays cheap because it reuses durable run data and a handful of bounded signals, not because it chases a particular service tier. More important, it tells the truth during failure and rollback. The web process can be green while imports are dead; the ledger, age rule, and stuck-run rule make that silence visible.

Sources

Top comments (0)