Decision rule: choose the dashboard path that can reconstruct one shipment-processing incident across product events and backend counters without weakening EU data controls. For a nightly logistics pipeline, that usually means keeping raw events and counters in separate, purpose-built stores, then joining small, stable aggregates through a thin read API. Plausible, PostHog, and Grafana Cloud can each occupy part of that design; a custom API can join it, but should not quietly become a fourth telemetry database.
Short answer: do not select on the prettiest chart or the lowest advertised tier. Test whether the system preserves event time, metric meaning, tenant boundaries, and a verifiable path from a KPI anomaly to the affected pipeline run. If it cannot answer which run changed which KPI, the dashboard is decoration during an incident.
This is an architecture decision record for a logistics SaaS whose nightly pipeline imports shipment updates, resolves status transitions, and publishes customer-facing summaries. The product view needs KPIs such as completed searches and successful exports. The operational view needs counters such as records read, records rejected, retries, and output rows. Incident reconstruction is the primary axis because an apparent conversion drop may be a product change, late input, a failed transform, or merely a disagreement between timestamps.
What must remain true?
The first invariant is semantic: every displayed number has one owner, one unit, and one documented aggregation rule. A shipments_rejected counter is not interchangeable with a product event called shipment_import_failed; one may count rows and the other user actions. Giving them similar labels does not make them comparable.
The second invariant is temporal. Each pipeline fact needs an event timestamp and a run identifier, while ingestion time remains separate. A file arriving after midnight can belong to the prior business day, and rewriting that history around arrival time makes a clean graph that tells the wrong story. Store both clocks.
The third invariant is privacy scope. Product events should carry the least identifying data needed for the stated measurement, backend metrics should avoid shipment or customer identifiers as labels, and the join layer should work from coarse tenant or region keys only when those keys are actually required. This is partly a legal question, so engineering sign-off is not enough; the data controller must establish purpose, retention, and access policy.
The fourth invariant is durability at the boundary. A dashboard refresh may fail without losing facts. A telemetry write may be retried without double-counting. A correction to a completed run must be represented explicitly rather than overwriting the evidence that an investigator needs later.
Counts disagree.
These constraints are intentionally stricter than “can draw a line chart.” They also expose the limits. Metrics are good at trends and alerts, event analytics is good at behavior slices, and neither automatically supplies a forensic record of every row-level transformation.
Where does reconstruction fail?
The common failure is cardinality disguised as convenience. Putting shipment_id, file_name, or an unbounded error string into metric labels seems useful for one investigation, then turns the metric space into a shadow event store. Keep bounded dimensions such as pipeline_stage, region, and result; put high-cardinality evidence in structured logs or a governed audit dataset.
A second failure is sampling the evidence that proves completeness. OpenTelemetry distinguishes head sampling, decided before a trace is complete, from tail sampling, decided after all or most spans are available. Either policy can be reasonable for traces, but a sampled trace stream must not become the authoritative denominator for records processed. The pipeline counter and run manifest need their own completeness contract.
Sampling is lossy.
Late and duplicate delivery form another boundary. At-least-once producers can emit the same update again, so the aggregator needs a stable idempotency key. Out-of-order updates require a declared event-time window and a visible correction policy. Without those rules, two dashboard backends can ingest the same source and honestly report different daily totals.
Then there is partial success. Imagine a run that reads 8,412,906 rows, rejects 317 malformed records, retries two output partitions, and publishes the other partitions. “Run succeeded” throws away the useful shape of the result. “Run failed” does too. The manifest should record stage outcomes and counts, with publication status separate from transform status. Those numbers are illustrative schema examples, not benchmark results.
One more trap matters in EU deployments: a region setting is not a privacy architecture. Collection fields, subprocessors, access paths, deletion behavior, backups, and export controls still require review. No dashboard label can settle those obligations.
Location alone proves little.
How should an EU SaaS metrics dashboard compare product KPIs and backend counters?
The products in the original shortlist are not substitutes at every layer. They expose different natural centers of gravity, and a responsible evaluation should verify current behavior in each product's own documentation and contract before procurement. The useful comparison is therefore not a ranking; it is a boundary test.
| Option | Natural role in this record | Evidence to demand in a proof of concept | Failure boundary to keep explicit |
|---|---|---|---|
| Plausible | Product-level aggregate input | Can the required KPI dimensions, retention, export, deletion, and EU processing terms be demonstrated with representative data? | Do not assume aggregate web or product events equal backend work completed. |
| PostHog | Product-event input and behavioral exploration | Can event-time corrections, identity policy, retention, access control, export, and deletion be tested against the governance plan? | Do not let flexible event properties become an undocumented warehouse schema. |
| Grafana Cloud | Operational metrics and investigation input | Can bounded counter labels, query semantics, retention, access control, export, and EU processing terms meet the runbook? | Do not encode per-shipment evidence as metric labels or treat sampled traces as totals. |
| Custom metrics API | Read-time normalization across governed aggregates | Can it reproduce a historical answer, expose provenance, enforce authorization, and survive a downstream outage? | The team owns schema evolution, deduplication, caching, on-call load, and every misleading join. |
This table deliberately avoids feature-checkmark theater. Product behavior and service terms can change, while the evidence demanded by the architecture should remain stable. Run the same acceptance dataset through every candidate: one ordinary run, one duplicate batch, one late batch, one partial publication, and one corrected run. Compare the returned values and provenance, not screenshots.
Cost belongs in the review, but as a consequence of shape: event volume, active series, query frequency, retention, export volume, and staff time. A cheap ingestion path can be expensive to investigate if it erases run identity; a custom endpoint can have a small infrastructure bill while imposing a large operational obligation. Ask each option for the same workload model and failure drill.
Limitations and trade-offs are part of the result. An aggregate-first product is not suitable when an investigator needs raw behavioral sequences; a flexible event system is a poor fit when the team cannot govern property growth and identity; a metrics service cannot replace a row-level audit trail; and a custom API is the wrong choice when nobody owns schema migrations, authorization, and on-call support. Choose the boundary the team can operate, then keep authoritative logistics records outside the dashboard layer.
The critical read path
The join API should accept a bounded time range and return aggregates with provenance. It should not accept arbitrary metric names, raw query fragments, or user identifiers. The following Python sketch shows the contract and the checks that matter; product_reader and counter_reader are generic adapters, not claims about any vendor API.
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from typing import Protocol
@dataclass(frozen=True)
class Window:
start: datetime
end: datetime
class AggregateReader(Protocol):
def read(self, *, tenant_scope: str, window: Window) -> dict: ...
def incident_snapshot(
tenant_scope: str,
window: Window,
product_reader: AggregateReader,
counter_reader: AggregateReader,
) -> dict:
if window.start.tzinfo is None or window.end.tzinfo is None:
raise ValueError("window must be timezone-aware")
if window.end <= window.start:
raise ValueError("window end must follow start")
if window.end - window.start > timedelta(days=31):
raise ValueError("incident window exceeds 31 days")
product = product_reader.read(
tenant_scope=tenant_scope, window=window
)
counters = counter_reader.read(
tenant_scope=tenant_scope, window=window
)
return {
"window": {
"start": window.start.astimezone(timezone.utc).isoformat(),
"end": window.end.astimezone(timezone.utc).isoformat(),
},
"product_kpis": product["values"],
"pipeline_counters": counters["values"],
"provenance": [product["provenance"], counters["provenance"]],
}
A production implementation also needs authorization tied to tenant_scope, timeouts, bounded concurrency, stale-cache behavior, and an error model that distinguishes “zero” from “source unavailable.” Those are correctness requirements. Returning zero after a timed-out upstream call is a particularly destructive shortcut because it converts missing evidence into a business event.
Zero is data.
Deployment should be boring: version the response schema, replay the fixed acceptance dataset before release, and compare old and new responses. Monitor adapter latency, error count, cache age, and provenance gaps. Keep the dashboard available from the last valid snapshot during a temporary source outage, but mark its age plainly; never present cached data as current.
Why reject one unified event stream?
A single stream for clicks, shipment states, pipeline counters, traces, and audit evidence is attractive because every chart appears to share one query model. It also concentrates incompatible retention, cardinality, access, and correction rules in one schema. For this decision, that is the rejected option. The blast radius is too wide, and the apparent simplicity moves complexity into conventions that are hard to enforce.
It still has a valid use case. A small system with one bounded event vocabulary, one access class, modest volume, and no need to distinguish sampled diagnostic data from authoritative counts may reasonably begin with one store. The exit criteria should be written before adoption: unbounded dimensions, conflicting retention duties, repeated backfills, or incident queries that interfere with ingestion are signals to separate workloads.
The final choice can be a product analytics input, an operational metrics input, and a narrow join API, or fewer components if the acceptance dataset proves they are unnecessary. The decisive result is reconstructability: for any disputed night, an operator can identify the run, the two clocks, the aggregate definitions, source availability, and any correction without treating a dashboard vendor as the system of record.
References
- OpenTelemetry, “Sampling”: https://opentelemetry.io/docs/concepts/sampling/
- Electron, “crashReporter”: https://www.electronjs.org/docs/latest/api/crash-reporter
Top comments (0)