A newly installed analytics stack can show a convincing sequence of "app open → product view → add to cart → purchase" before it has ingested a single real user event. That is useful for testing whether the database and dashboard start, but it is not evidence of conversion. Four non-empty bars may represent four disjoint groups of users.
This is a practical audit of SensorFlow's public demo-data generator. SensorFlow is an independent, open-source, self-hosted event pipeline built around an existing Sensors Data SDK integration, a Go receiver, ClickHouse, and Apache Superset. It does not claim to replace every feature of a full product-analytics suite. Its installation deliberately shows synthetic data first; live SDK ingestion requires a separate license activation. That boundary makes the demo a useful example of why a chart needs a data-contract check.
Does a four-stage dashboard prove a conversion funnel?
No. Stage counts answer "did these event names occur?" A user funnel answers "did the same identified user perform these events in the required order, within the chosen time window?" Counts, distinct-user counts, user overlap, and ordered conversion are four different measures.
For a real funnel, define the identity key, the start and end events, a time window, timezone, duplicate handling, and whether anonymous and logged-in IDs are merged. Without those decisions, a percentage can look precise while measuring the wrong thing.
Inspect the generator before interpreting the chart
The current demo SQL inserts rows from numbers(240). It selects the event name with number % 12, and the demo user with number % 48. Because 48 is divisible by 12, each demo-user-* stays in one event category throughout this generated batch. A demo user is not walking through four stages.
The script also spaces event timestamps three hours apart and assigns sample revenue to purchase rows. Its is_first_day value is part of the synthetic generator, not a production calculation after identity stitching. These fields are good for rendering example charts, not for validating business conversion or GMV.
The underlying ClickHouse table is sensors.event, with time, event, distinct_id, and app_id. The demo rows are marked app_id = 'sensorflow-demo'. If you have installed the bundled demo, first compare event rows with distinct users:
SELECT
event,
count() AS event_rows,
uniqExact(distinct_id) AS users
FROM sensors.event
WHERE app_id = 'sensorflow-demo'
GROUP BY event
ORDER BY event;
On a fresh import of one batch from the current generator, this returns 100 app-open rows from 20 IDs, 80 product-view rows from 16 IDs, 40 add-to-cart rows from 8 IDs, and 20 purchase rows from 4 IDs. Those numbers come from the generator, not from measured customer behavior. Existing installations or changed scripts may produce different results.
Test whether the viewers and buyers are even the same people
Before reporting 4 / 16 = 25% as a conversion rate, ask whether a single distinct_id appears in both groups:
WITH per_user AS
(
SELECT
distinct_id,
countIf(event = 'demo_product_view') > 0 AS viewed,
countIf(event = 'demo_purchase') > 0 AS purchased
FROM sensors.event
WHERE app_id = 'sensorflow-demo'
GROUP BY distinct_id
)
SELECT
countIf(viewed) AS view_users,
countIf(purchased) AS purchase_users,
countIf(viewed AND purchased) AS both_users
FROM per_user;
Recreating the published generator's 240 rows with ClickHouse Local produced 16 / 4 / 0: sixteen viewers, four buyers, and zero IDs in both sets. Therefore 4 / 16 is not a view-to-purchase conversion rate for this demo. Even a non-zero overlap would establish only that the events occurred for the same ID, not that they happened in the right order.
The query uses ClickHouse's documented countIf aggregate combinator and uniqExact. Exact distinct counts are useful when checking a small test set; ClickHouse notes that uniqExact can consume more memory than approximate alternatives as cardinality grows. Do not infer production query performance from this 240-row sample.
What a production funnel must add
First, choose a stable identity rule. Decide whether distinct_id represents a logged-in account, an anonymous browser ID, or a merged identity. Cross-device behavior and login transitions can change the denominator.
Second, enforce sequence and time. A user who purchases before the product-view event must not count as a view-to-purchase conversion under an ordered definition. ClickHouse documents windowFunnel for ordered chains inside a sliding window. Its window units depend on the timestamp argument; SensorFlow's time column is DateTime64(3), so do not paste a seconds-based example into a millisecond conversion without testing it against your ClickHouse version.
Third, isolate environments and reconcile the dashboard. Filter out sensorflow-demo, staging traffic, and test accounts when calculating production KPIs. A Superset bar chart should use the same event names, identity rule, date boundaries, and filters as a SQL check against raw events.
Finally, test the real pipeline with a small, manually checkable set of SDK events. Confirm the client request, receiver response, ClickHouse row, and Superset metric separately. SensorFlow's SDK integration notes describe the receiver URL and compatibility boundaries. The README makes clear that installing the demo does not start live ingestion; activating it is a separate licensed step.
A five-step acceptance test
- Read the demo or seed-data generator; determine whether it represents a real user journey at all.
- Compare raw rows and distinct IDs for every event stage.
- Check overlap between stages before calculating any conversion percentage.
- Add ordered-event, time-window, identity, and deduplication rules to the real-event test.
- Match the dashboard filters to the SQL query and exclude synthetic traffic from production reports.
The takeaway is not that synthetic dashboards are bad. They are excellent smoke tests for installation and visualization. The mistake is treating their stage distribution as evidence that real users converted. SensorFlow's transparent ClickHouse path lets a team inspect raw events and define its own metrics, but it also leaves that metric design and infrastructure operation with the team. If your organization needs a ready-made, no-SQL product-analytics interface instead, evaluate one on that basis.
Top comments (0)