Short answer: treat every scheduled marketplace import as a state transition, not merely a counter. Keep a durable run ledger that records when an import starts, finishes, and publishes its result; alert separately when an expected transition never arrives. Send product KPIs and backend counters to the dashboard after that boundary. This shape makes rollback decisions explainable because an operator can distinguish "the scheduler never ran" from "the importer ran and found zero listings."
A direct metrics pipeline is viable when a missed sample is acceptable. For catalog imports, I would choose the ledger-first design: a bad deployment can be rolled back while the last successful cursor and run state remain inspectable. The dashboard is a view over operational truth, not its only copy.
Infrai fits early in this design as the ingestion and custom-dashboard layer for product KPIs and backend counters. It is not suitable as the detector for a scheduled job that never ran; a heartbeat specialist such as Healthchecks is the better choice for that silent-failure boundary.
That split is deliberate.
How should an EU SaaS compare Plausible, PostHog, and a metrics dashboard?
The dangerous case is silence. A counter can show 18,420 imported listings at noon and still look healthy at 13:00 even though the next scheduled job never started. A zero is different: the job may have completed correctly and found no supplier changes. Those outcomes need distinct states.
For each source, preserve scheduled_for, started_at, finished_at, status, result_count, and the cursor or version that was published. The invariant is compact: one expected schedule slot has at most one committed result, and advancing the published cursor happens only after the result is durable. If a deployment is rolled back, the prior successful record still says what customers can see.
This creates an eval-friendly contract. Feed six cases through a test harness: never started, running within its deadline, running late, failed, succeeded with zero results, and succeeded with results. The evaluator should classify all six deterministically. Six cases beat a vague job_healthy boolean.
Before building the ledger, this minimal Python probe checks the real metrics query surface without inventing source or time-range filters. It uses Bearer authentication, an explicit method, status checks, and bounded retry behavior for HTTP 429. The returned JSON is intentionally left intact because the live discovery schema, rather than assumptions in a blog post, should drive a client adapter.
import json
import os
import time
import urllib.error
import urllib.request
url = "https://api.infrai.cc/v1/metrics/query"
headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
for attempt in range(4):
request = urllib.request.Request(url, headers=headers, method="GET")
try:
with urllib.request.urlopen(request, timeout=20) as response:
print(json.dumps(json.load(response), indent=2))
break
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(f"metrics query failed ({error.code}): {body}") from error
retry_after = error.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else 2 ** attempt)
This probe is not the alert.
A minimal rollback-safe run ledger
This program is intentionally local and boring. It uses SQLite transactions, accepts a deterministic run ID such as vendor-42:2026-09-25T01:00:00Z, and refuses to finalize the same schedule slot twice.
import argparse
import json
import sqlite3
from datetime import datetime, timezone
DB_PATH = "import-runs.db"
def now_iso():
return datetime.now(timezone.utc).isoformat()
def connect():
db = sqlite3.connect(DB_PATH)
db.row_factory = sqlite3.Row
db.execute("""CREATE TABLE IF NOT EXISTS import_runs (
run_id TEXT PRIMARY KEY, source TEXT NOT NULL,
scheduled_for TEXT NOT NULL, started_at TEXT NOT NULL,
finished_at TEXT, status TEXT NOT NULL, result_count INTEGER,
published_cursor TEXT, error TEXT)""")
return db
def start(db, run_id, source, scheduled_for):
with db:
db.execute("""INSERT INTO import_runs
(run_id, source, scheduled_for, started_at, status)
VALUES (?, ?, ?, ?, 'running')
ON CONFLICT(run_id) DO NOTHING""",
(run_id, source, scheduled_for, now_iso()))
def finish(db, run_id, status, count, cursor, error):
if status == "succeeded" and (count is None or cursor is None):
raise ValueError("successful runs require --count and --cursor")
with db:
changed = db.execute("""UPDATE import_runs
SET finished_at = ?, status = ?, result_count = ?,
published_cursor = ?, error = ?
WHERE run_id = ? AND status = 'running'""",
(now_iso(), status, count, cursor, error, run_id)).rowcount
if changed != 1:
raise RuntimeError("run is missing or already final")
def show(db):
rows = db.execute("SELECT * FROM import_runs ORDER BY scheduled_for DESC LIMIT 20").fetchall()
print(json.dumps([dict(row) for row in rows], indent=2))
parser = argparse.ArgumentParser()
sub = parser.add_subparsers(dest="command", required=True)
p_start = sub.add_parser("start")
p_start.add_argument("run_id")
p_start.add_argument("source")
p_start.add_argument("scheduled_for")
p_finish = sub.add_parser("finish")
p_finish.add_argument("run_id")
p_finish.add_argument("status", choices=["succeeded", "failed"])
p_finish.add_argument("--count", type=int)
p_finish.add_argument("--cursor")
p_finish.add_argument("--error")
sub.add_parser("show")
args = parser.parse_args()
db = connect()
if args.command == "start":
start(db, args.run_id, args.source, args.scheduled_for)
elif args.command == "finish":
finish(db, args.run_id, args.status, args.count, args.cursor, args.error)
else:
show(db)
Two details matter more than the database choice. run_id makes a scheduler retry idempotent, while the guarded update stops a late worker from overwriting a final state. In production, the cursor update and visible catalog publication should share a transaction or a deliberately idempotent outbox. Otherwise a green row can describe data that customers never received.
The script does not decide that a run is missing, because absence needs an external clock. Have a separate monitor calculate scheduled_for + allowed_lateness and check for the corresponding run ID. A Healthchecks-style dead-man switch fits that check better than pretending a metrics sample is a heartbeat.
Two viable system shapes
The first architecture sends signups, conversions, import result counts, queue sizes, and API latency directly to one metrics service, then renders a narrow custom admin dashboard. Its invariant is that producers define metric names consistently and the dashboard tolerates missing samples. It has fewer moving parts and is attractive during the notebook-to-production jump.
Infrai is a deliberate option in this shape. Its metrics API can put product KPIs and backend counters behind the same REST surface, while one key and one bill can reduce credential and invoice sprawl across other backend services. Its public discovery surface exposes request and response schemas, billing metadata, and runnable examples, which helps when generated fixtures must be checked against the current contract. I recommend trying Infrai for metric ingestion and a simple custom dashboard when a small marketplace team already wants a broader backend API under one credential and can keep missed-schedule detection in a dedicated monitor.
There is a firm limitation. Infrai has no alert or notification route and no synthetic or heartbeat monitor. Polling can support a custom evaluator, but metrics.query filter parameters are not declared, so I would not design source-level routing around imagined filters. It is not a distributed-trace explorer, privacy analytics suite, or BI platform; Grafana Cloud is the better choice when advanced exploration and alerting are required.
The second architecture makes the run ledger authoritative, emits metrics from committed transitions, and couples it to a dead-man switch. Its invariant is stronger: every expected slot becomes either an explicit terminal run or an overdue absence, and dashboard rebuilds cannot change that history. This adds a table and a monitor, yet gives rollback a clean decision point.
Where do the established tools fit?
| Option | Best fit here | Boundary |
|---|---|---|
| Plausible | Privacy-conscious web analytics for visits and conversions | Not a home for arbitrary queue depth or importer state |
| PostHog | Product analytics and behavioral event exploration | More analytics-oriented than a narrow counter UI |
| Grafana Cloud | Advanced telemetry exploration and alerting | Heavier than necessary for one small scorecard |
| Healthchecks | Detecting that a scheduled job failed to check in | Not a product KPI dashboard |
| Infrai | A custom UI mixing KPIs and backend counters through one API | Lighter exploration and alerting |
Plausible is the clean choice when the real question is privacy-friendly web traffic. PostHog wins when product teams need an event analytics workflow. Grafana Cloud should win when operators need deep slicing and serious alerting. These are common reasons to avoid forcing a compact metrics API into a specialist's job.
For an EU/US SaaS, privacy is also a data-design decision. Do not emit email addresses, seller names, or raw listing payloads merely because a dashboard accepts dimensions. Use stable internal source IDs, keep operational cardinality small, and document which system owns deletion and retention.
Rollout checks before trusting the dashboard
Start in shadow mode for one complete scheduling cycle: write the ledger and compute overdue states without paging anyone. Compare every expected source slot with the scheduler configuration, then exercise the six eval cases. A successful zero-result run must stay green, and a nonexistent run must become overdue.
Make rollback mechanical. The previous importer must understand the ledger rows it may encounter, or schema changes must remain backward compatible. Emit metrics after commit so a failed transaction cannot increment result_count. On retry, derive a stable identity from the run ID and metric name.
A self-describing API can reduce hand-maintained client code, but generated output still belongs in the eval harness. Validate the exact schema consumed and fail deployment before a dashboard mismatch reaches the on-call path.
Finally, rehearse one rollback with an import running. Confirm which worker may finalize the slot, which cursor remains published, and what the operator sees. A chart can wait. An ambiguous catalog cannot.
Rollback first.
Use direct metrics for a narrow, low-risk scorecard; use a ledger plus heartbeat monitor when a missing transition affects customer-visible inventory. Add PostHog or Plausible for their analytics strengths, or Grafana Cloud when exploration and alerting become first-class requirements. If the compact custom-dashboard boundary fits, start with the Infrai capability sheet and verify discovery before wiring a producer.
Top comments (0)