<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: SensorFlow</title>
    <description>The latest articles on DEV Community by SensorFlow (@sensorflow).</description>
    <link>https://dev.to/sensorflow</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4123045%2Fd2cb1876-ab98-454b-b579-f51a9afe52b4.png</url>
      <title>DEV Community: SensorFlow</title>
      <link>https://dev.to/sensorflow</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sensorflow"/>
    <language>en</language>
    <item>
      <title>Design Events That Answer Product Questions: A Practical Tracking Plan</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Thu, 01 Oct 2026 06:43:59 +0000</pubDate>
      <link>https://dev.to/sensorflow/design-events-that-answer-product-questions-a-practical-tracking-plan-2995</link>
      <guid>https://dev.to/sensorflow/design-events-that-answer-product-questions-a-practical-tracking-plan-2995</guid>
      <description>&lt;p&gt;When someone asks why signup conversion dropped, a generic &lt;code&gt;button_click&lt;/code&gt; event rarely answers the question. The team needs to agree on where the journey starts, what counts as completion, how users are counted, and how failures and retries are represented.&lt;/p&gt;

&lt;p&gt;An event tracking plan records those agreements before instrumentation. It connects a business question to event names, trigger conditions, identity rules, properties and acceptance tests. A working collector cannot resolve a disagreement about what an event means.&lt;/p&gt;

&lt;p&gt;This guide uses a fictional SaaS signup flow. The examples are teaching fixtures, not customer results. I work on SensorFlow, a self-hosted event analytics project; the design method below applies independently of the analytics tool you choose. AI assisted the preparation of this article, and its references and product claims were checked before publication.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with a question you can calculate
&lt;/h2&gt;

&lt;p&gt;Instead of asking for "signup conversion," write a more precise question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Of the distinct visitors who opened the signup page during a calendar day in our reporting timezone, how many completed account creation within 24 hours of their first qualifying page visit?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This defines the starting population, completion condition, identity unit, reporting timezone and conversion window. You still need to define how an anonymous visitor maps to an account, and whether repeated visits restart the window. Those decisions belong in the metric definition.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://amplitude.com/docs/get-started/select-events" rel="noopener noreferrer"&gt;Amplitude's event selection guide&lt;/a&gt; works back from the questions a product team wants to answer. Capturing more interactions does not automatically produce a better measurement plan. Autocapture can help explore interface behavior, but a click or form submission is not proof that an account was created. &lt;a href="https://posthog.com/docs/product-analytics/capture-events" rel="noopener noreferrer"&gt;PostHog documents automatic capture and custom events separately&lt;/a&gt;; custom business events let you express the outcome your system can actually confirm.&lt;/p&gt;

&lt;h2&gt;
  
  
  Define a small event contract
&lt;/h2&gt;

&lt;p&gt;For each event, write down the following fields. The table describes a suggested contract, rather than a required schema for any vendor.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Question to resolve&lt;/th&gt;
&lt;th&gt;Signup example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Event name&lt;/td&gt;
&lt;td&gt;What action or state occurred?&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;signup_submitted&lt;/code&gt;, &lt;code&gt;account_created&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trigger&lt;/td&gt;
&lt;td&gt;At what exact point is it emitted?&lt;/td&gt;
&lt;td&gt;Emit &lt;code&gt;account_created&lt;/code&gt; after the account transaction commits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Source&lt;/td&gt;
&lt;td&gt;Which component owns that state?&lt;/td&gt;
&lt;td&gt;Browser for intent; backend for confirmed creation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Identity&lt;/td&gt;
&lt;td&gt;Which stable identifier is used?&lt;/td&gt;
&lt;td&gt;Anonymous visitor before signup; account ID after creation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time&lt;/td&gt;
&lt;td&gt;When did it occur, and which timezone does the report use?&lt;/td&gt;
&lt;td&gt;Preserve occurrence time and define the reporting timezone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Required properties&lt;/td&gt;
&lt;td&gt;Which dimensions are essential?&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;signup_method&lt;/code&gt;, &lt;code&gt;app_version&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Optional properties&lt;/td&gt;
&lt;td&gt;How is missing context represented?&lt;/td&gt;
&lt;td&gt;Missing &lt;code&gt;campaign_id&lt;/code&gt; is unknown, not silently "direct"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retry rule&lt;/td&gt;
&lt;td&gt;Can one operation produce multiple events?&lt;/td&gt;
&lt;td&gt;Reuse a stable operation ID and define deduplication&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Owner and version&lt;/td&gt;
&lt;td&gt;Who approves a changed meaning?&lt;/td&gt;
&lt;td&gt;Product and engineering owners, plus a contract version&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://docs.snowplow.io/docs/fundamentals/tracking-design-best-practice/" rel="noopener noreferrer"&gt;Snowplow's tracking design documentation&lt;/a&gt; combines event purpose, trigger conditions and data structures in event specifications. &lt;a href="https://amplitude.com/docs/data/create-tracking-plan" rel="noopener noreferrer"&gt;Amplitude's tracking plan documentation&lt;/a&gt; also records sources, descriptions and property rules. Both are useful references when turning a short product request into something an engineer can implement and a tester can verify.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep names stable and put context in properties
&lt;/h2&gt;

&lt;p&gt;Avoid separate event names such as &lt;code&gt;signup_from_ad&lt;/code&gt; and &lt;code&gt;signup_from_email&lt;/code&gt; when the action is the same. A stable &lt;code&gt;signup_submitted&lt;/code&gt; event with a &lt;code&gt;signup_source&lt;/code&gt; property is easier to extend when another channel appears. &lt;a href="https://www.twilio.com/docs/segment/protocols/tracking-plan/best-practices.md" rel="noopener noreferrer"&gt;Segment's tracking plan guidance&lt;/a&gt; similarly advises against dynamic event names and property keys.&lt;/p&gt;

&lt;p&gt;Choose a naming convention and apply it consistently. Title Case or snake_case can both work; the problem is a team alternating between them without a defined mapping. Give each property a type, allowed values, and a rule for missing data. Keep email addresses, phone numbers and full form contents out of a generic analytics payload unless a separately reviewed use case requires them.&lt;/p&gt;

&lt;p&gt;Do not merge intent and outcome just to reduce the event count. &lt;code&gt;signup_submitted&lt;/code&gt; means a user tried. &lt;code&gt;account_created&lt;/code&gt; means the business state changed. If both become a vague &lt;code&gt;signup&lt;/code&gt;, a validation error can look like a successful conversion.&lt;/p&gt;

&lt;p&gt;Some tools provide recommended events, such as those described in &lt;a href="https://developers.google.com/analytics/devguides/collection/ga4/events" rel="noopener noreferrer"&gt;Google Analytics' event guide&lt;/a&gt;. Define your own business meaning first, then map it to each destination. Similar event names do not guarantee equivalent semantics across tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instrument at the agreed trigger
&lt;/h2&gt;

&lt;p&gt;For a project already using the official Sensors Data JavaScript SDK, a custom intent event could have this shape after SDK initialization:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Record the attempt. This is not proof of account creation.&lt;/span&gt;
&lt;span class="nx"&gt;sensors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;track&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;signup_submitted&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;signup_method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;email&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;signup_source&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;website&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;contract_version&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;1&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is an instrumentation fragment, not a standalone application or a live ingestion test. Initialize the SDK, identity handling and destination for your own environment; &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/docs/sensors-sdk.zh-CN.md" rel="noopener noreferrer"&gt;SensorFlow's SDK integration guide&lt;/a&gt; provides project-specific setup references. Consult the official documentation for your exact SDK version.&lt;/p&gt;

&lt;p&gt;Let the component that can confirm the final business state own the success event. An HTTP 200 response might only mean a request was accepted. If the browser and backend both emit &lt;code&gt;account_created&lt;/code&gt;, define how they are reconciled, and where a stable operation ID is generated. A property called &lt;code&gt;operation_id&lt;/code&gt; does not create deduplication by itself: the receiving and analytical layers must use it according to the contract.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the contract, then test the metric
&lt;/h2&gt;

&lt;p&gt;Prepare a small set of non-sensitive fixtures with expected event sequences:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;An anonymous visitor opens the signup page and leaves. Expect the start event and no creation event.&lt;/li&gt;
&lt;li&gt;Form validation fails. Specify whether &lt;code&gt;signup_submitted&lt;/code&gt; means every submit attempt or only a valid request; expect no creation event in either case.&lt;/li&gt;
&lt;li&gt;An account is created. Expect one confirmed creation with the correct property types.&lt;/li&gt;
&lt;li&gt;A network failure causes a retry. Check how the same operation is counted, not just how many requests were sent.&lt;/li&gt;
&lt;li&gt;A visitor signs in or continues on another device. Verify the documented anonymous-to-account mapping instead of assuming two tools merge identities identically.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Check raw records against the contract first. Then recalculate the metric using the same identity, timezone, window and retry rules as the report. Receiving an event and calculating an interpretable conversion rate are separate acceptance criteria.&lt;/p&gt;

&lt;p&gt;Keep expected and actual results with the release record. Those fixtures become regression tests when the SDK, application workflow or analytics destination changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose an analytics system against the contract
&lt;/h2&gt;

&lt;p&gt;An open-source or self-hosted analytics system should be evaluated against the events and questions your team needs. Check whether it receives your existing instrumentation, exposes raw records, explains identity and time behavior, and provides the analyses your users need. Also account for the deployment, access control, monitoring and upgrade work your team will own.&lt;/p&gt;

&lt;p&gt;SensorFlow is one option for teams retaining an existing Sensors Data SDK while controlling the event data pipeline. The open-source repository supports deployment and demonstration data; real SDK ingestion requires license activation. It is a focused self-hosted path rather than a feature-equivalent replacement for a mature all-in-one analytics suite, and it has no official affiliation with Sensors Data.&lt;/p&gt;

&lt;p&gt;Start by writing one event contract and checking these five fixtures. Once the team can explain the result by hand, it has a concrete basis for selecting an analytics system and reviewing its reports.&lt;/p&gt;

</description>
      <category>analytics</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Your Demo Dashboard Is Not a Funnel: Auditing Synthetic Events in ClickHouse</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Thu, 01 Oct 2026 04:49:25 +0000</pubDate>
      <link>https://dev.to/sensorflow/your-demo-dashboard-is-not-a-funnel-auditing-synthetic-events-in-clickhouse-527f</link>
      <guid>https://dev.to/sensorflow/your-demo-dashboard-is-not-a-funnel-auditing-synthetic-events-in-clickhouse-527f</guid>
      <description>&lt;p&gt;A newly installed analytics stack can show a convincing sequence of "app open → product view → add to cart → purchase" before it has ingested a single real user event. That is useful for testing whether the database and dashboard start, but it is not evidence of conversion. Four non-empty bars may represent four disjoint groups of users.&lt;/p&gt;

&lt;p&gt;This is a practical audit of &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/deploy/docker/clickhouse/demo_data.sql" rel="noopener noreferrer"&gt;SensorFlow's public demo-data generator&lt;/a&gt;. SensorFlow is an independent, open-source, self-hosted event pipeline built around an existing Sensors Data SDK integration, a Go receiver, ClickHouse, and Apache Superset. It does not claim to replace every feature of a full product-analytics suite. Its installation deliberately shows synthetic data first; live SDK ingestion requires a separate license activation. That boundary makes the demo a useful example of why a chart needs a data-contract check.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does a four-stage dashboard prove a conversion funnel?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;No.&lt;/strong&gt; Stage counts answer "did these event names occur?" A user funnel answers "did the same identified user perform these events in the required order, within the chosen time window?" Counts, distinct-user counts, user overlap, and ordered conversion are four different measures.&lt;/p&gt;

&lt;p&gt;For a real funnel, define the identity key, the start and end events, a time window, timezone, duplicate handling, and whether anonymous and logged-in IDs are merged. Without those decisions, a percentage can look precise while measuring the wrong thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Inspect the generator before interpreting the chart
&lt;/h2&gt;

&lt;p&gt;The current &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/deploy/docker/clickhouse/demo_data.sql" rel="noopener noreferrer"&gt;demo SQL&lt;/a&gt; inserts rows from &lt;code&gt;numbers(240)&lt;/code&gt;. It selects the event name with &lt;code&gt;number % 12&lt;/code&gt;, and the demo user with &lt;code&gt;number % 48&lt;/code&gt;. Because 48 is divisible by 12, each &lt;code&gt;demo-user-*&lt;/code&gt; stays in one event category throughout this generated batch. A demo user is not walking through four stages.&lt;/p&gt;

&lt;p&gt;The script also spaces event timestamps three hours apart and assigns sample revenue to purchase rows. Its &lt;code&gt;is_first_day&lt;/code&gt; value is part of the synthetic generator, not a production calculation after identity stitching. These fields are good for rendering example charts, not for validating business conversion or GMV.&lt;/p&gt;

&lt;p&gt;The underlying &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/deploy/docker/clickhouse/init_clickhouse.sql" rel="noopener noreferrer"&gt;ClickHouse table&lt;/a&gt; is &lt;code&gt;sensors.event&lt;/code&gt;, with &lt;code&gt;time&lt;/code&gt;, &lt;code&gt;event&lt;/code&gt;, &lt;code&gt;distinct_id&lt;/code&gt;, and &lt;code&gt;app_id&lt;/code&gt;. The demo rows are marked &lt;code&gt;app_id = 'sensorflow-demo'&lt;/code&gt;. If you have installed the bundled demo, first compare event rows with distinct users:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;event_rows&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;uniqExact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;distinct_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;sensors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;app_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'sensorflow-demo'&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a fresh import of one batch from the current generator, this returns 100 app-open rows from 20 IDs, 80 product-view rows from 16 IDs, 40 add-to-cart rows from 8 IDs, and 20 purchase rows from 4 IDs. Those numbers come from the generator, not from measured customer behavior. Existing installations or changed scripts may produce different results.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test whether the viewers and buyers are even the same people
&lt;/h2&gt;

&lt;p&gt;Before reporting &lt;code&gt;4 / 16 = 25%&lt;/code&gt; as a conversion rate, ask whether a single &lt;code&gt;distinct_id&lt;/code&gt; appears in both groups:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="n"&gt;per_user&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt;
&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt;
        &lt;span class="n"&gt;distinct_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;countIf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'demo_product_view'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;viewed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;countIf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'demo_purchase'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;purchased&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;sensors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;
    &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;app_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'sensorflow-demo'&lt;/span&gt;
    &lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;distinct_id&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;countIf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;viewed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;view_users&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;countIf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;purchased&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;purchase_users&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;countIf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;viewed&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;purchased&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;both_users&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;per_user&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Recreating the published generator's 240 rows with ClickHouse Local produced &lt;code&gt;16 / 4 / 0&lt;/code&gt;: sixteen viewers, four buyers, and zero IDs in both sets. Therefore &lt;code&gt;4 / 16&lt;/code&gt; is &lt;strong&gt;not&lt;/strong&gt; a view-to-purchase conversion rate for this demo. Even a non-zero overlap would establish only that the events occurred for the same ID, not that they happened in the right order.&lt;/p&gt;

&lt;p&gt;The query uses ClickHouse's documented &lt;a href="https://clickhouse.com/docs/reference/functions/aggregate-functions/combinators" rel="noopener noreferrer"&gt;&lt;code&gt;countIf&lt;/code&gt; aggregate combinator&lt;/a&gt; and &lt;a href="https://clickhouse.com/docs/reference/functions/aggregate-functions/uniqExact" rel="noopener noreferrer"&gt;&lt;code&gt;uniqExact&lt;/code&gt;&lt;/a&gt;. Exact distinct counts are useful when checking a small test set; ClickHouse notes that &lt;code&gt;uniqExact&lt;/code&gt; can consume more memory than approximate alternatives as cardinality grows. Do not infer production query performance from this 240-row sample.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a production funnel must add
&lt;/h2&gt;

&lt;p&gt;First, choose a stable identity rule. Decide whether &lt;code&gt;distinct_id&lt;/code&gt; represents a logged-in account, an anonymous browser ID, or a merged identity. Cross-device behavior and login transitions can change the denominator.&lt;/p&gt;

&lt;p&gt;Second, enforce sequence and time. A user who purchases before the product-view event must not count as a view-to-purchase conversion under an ordered definition. ClickHouse documents &lt;a href="https://clickhouse.com/docs/reference/functions/aggregate-functions/parametric-functions#windowfunnel" rel="noopener noreferrer"&gt;&lt;code&gt;windowFunnel&lt;/code&gt;&lt;/a&gt; for ordered chains inside a sliding window. Its window units depend on the timestamp argument; SensorFlow's &lt;code&gt;time&lt;/code&gt; column is &lt;code&gt;DateTime64(3)&lt;/code&gt;, so do not paste a seconds-based example into a millisecond conversion without testing it against your ClickHouse version.&lt;/p&gt;

&lt;p&gt;Third, isolate environments and reconcile the dashboard. Filter out &lt;code&gt;sensorflow-demo&lt;/code&gt;, staging traffic, and test accounts when calculating production KPIs. A Superset bar chart should use the same event names, identity rule, date boundaries, and filters as a SQL check against raw events.&lt;/p&gt;

&lt;p&gt;Finally, test the real pipeline with a small, manually checkable set of SDK events. Confirm the client request, receiver response, ClickHouse row, and Superset metric separately. &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/docs/sensors-sdk.zh-CN.md" rel="noopener noreferrer"&gt;SensorFlow's SDK integration notes&lt;/a&gt; describe the receiver URL and compatibility boundaries. &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/README.md" rel="noopener noreferrer"&gt;The README&lt;/a&gt; makes clear that installing the demo does not start live ingestion; activating it is a separate licensed step.&lt;/p&gt;

&lt;h2&gt;
  
  
  A five-step acceptance test
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Read the demo or seed-data generator; determine whether it represents a real user journey at all.&lt;/li&gt;
&lt;li&gt;Compare raw rows and distinct IDs for every event stage.&lt;/li&gt;
&lt;li&gt;Check overlap between stages before calculating any conversion percentage.&lt;/li&gt;
&lt;li&gt;Add ordered-event, time-window, identity, and deduplication rules to the real-event test.&lt;/li&gt;
&lt;li&gt;Match the dashboard filters to the SQL query and exclude synthetic traffic from production reports.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The takeaway is not that synthetic dashboards are bad. They are excellent smoke tests for installation and visualization. The mistake is treating their stage distribution as evidence that real users converted. SensorFlow's transparent ClickHouse path lets a team inspect raw events and define its own metrics, but it also leaves that metric design and infrastructure operation with the team. If your organization needs a ready-made, no-SQL product-analytics interface instead, evaluate one on that basis.&lt;/p&gt;

</description>
      <category>analytics</category>
      <category>database</category>
      <category>sql</category>
      <category>opensource</category>
    </item>
    <item>
      <title>SensorFlow vs Countly in 2026: Analytics Application or SDK-to-ClickHouse Pipeline?</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Tue, 29 Sep 2026 08:00:00 +0000</pubDate>
      <link>https://dev.to/sensorflow/sensorflow-vs-countly-in-2026-analytics-application-or-sdk-to-clickhouse-pipeline-3o9m</link>
      <guid>https://dev.to/sensorflow/sensorflow-vs-countly-in-2026-analytics-application-or-sdk-to-clickhouse-pipeline-3o9m</guid>
      <description>&lt;p&gt;Countly and SensorFlow both matter to teams that want to operate analytics infrastructure, but they solve different starting problems. Countly is an analytics application with its own SDKs, event collection, reporting interface, and engagement features. SensorFlow is a narrower ingestion path for compatible existing Sensors Data SDK events: Go receiver → team-owned ClickHouse → SQL or Apache Superset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; Start with Countly if you need an integrated web, mobile, or desktop analytics product and can instrument with Countly SDKs. Test SensorFlow if you already use Sensors Data SDK, want to keep that instrumentation, and specifically need events in your own ClickHouse. Neither choice removes the need to check edition terms, actual payload compatibility, privacy duties, and operational cost.&lt;/p&gt;

&lt;p&gt;This comparison uses the projects' public documentation checked in September 2026. It is not a benchmark, legal opinion, or price quote. SensorFlow is independent of Countly and Sensors Data.&lt;/p&gt;

&lt;h2&gt;
  
  
  First decide what you are replacing
&lt;/h2&gt;

&lt;p&gt;Countly's &lt;a href="https://github.com/Countly/countly-server/blob/master/README.md" rel="noopener noreferrer"&gt;public server README&lt;/a&gt; describes mobile, web, and desktop collection, sessions, views, events, dashboards, APIs, and plugins. Countly supplies its own SDKs across Lite, Flex, and Enterprise. A team adopting it can instrument an app with a Countly SDK and use Countly's reporting interface. It is not merely a pageview counter.&lt;/p&gt;

&lt;p&gt;SensorFlow starts from another constraint: an application already sends events through a compatible Sensors Data SDK and the team does not want to rewrite every client just to change where data lands. The &lt;a href="https://github.com/data-analyze-bi/sensorFlow" rel="noopener noreferrer"&gt;SensorFlow repository&lt;/a&gt; documents this route:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Existing Sensors Data SDK clients send requests to a Go ingestion service.&lt;/li&gt;
&lt;li&gt;The service writes event rows to the team's ClickHouse.&lt;/li&gt;
&lt;li&gt;Analysts query ClickHouse directly or build Apache Superset dashboards.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The repository offers a license-free demo with sample data and Superset dashboards. &lt;strong&gt;Real SDK ingestion requires a SensorFlow license.&lt;/strong&gt; SensorFlow is not a replacement for Countly's built-in analytics and engagement UI, nor a promise that every Sensors Data SDK version or payload mode works automatically.&lt;/p&gt;

&lt;p&gt;Countly's &lt;a href="https://api.count.ly/reference" rel="noopener noreferrer"&gt;write API&lt;/a&gt; can accept data from other sources, but that is not the same as accepting Sensors Data SDK wire requests. Pointing an existing SDK server URL at Countly's collector is not a demonstrated migration; it would require an adapter or re-instrumentation and identity testing. Countly SDK calls likewise are not Sensors Data SDK calls that SensorFlow can ingest as-is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compare the edition you would actually run
&lt;/h2&gt;

&lt;p&gt;Countly Lite is self-hosted and includes core event collection and dashboards. Countly's current &lt;a href="https://countly.com/features" rel="noopener noreferrer"&gt;feature matrix&lt;/a&gt; lists dashboards, event management, and mobile/web/desktop analytics for Lite. The same matrix marks funnels, retention, cohorts, and user profiles for Flex/Enterprise, not Lite. It would be misleading either to say “Countly has no funnels” or to promise that Lite includes them. Check the exact version and plan you are evaluating; packaging can change.&lt;/p&gt;

&lt;p&gt;Countly's &lt;a href="https://github.com/Countly/countly-server/blob/master/LICENSE.md" rel="noopener noreferrer"&gt;server license&lt;/a&gt; states AGPL-3.0 with modified Section 7 and branding restrictions. The &lt;a href="https://support.countly.com/hc/en-us/articles/360037501312-Countly-Licensing-FAQ" rel="noopener noreferrer"&gt;official licensing FAQ&lt;/a&gt; explains an important nuance: organizations can use Lite internally and track their own paid apps, but offering Countly as a service to customers, giving those customers dashboard access for their apps, or rebranding it has different licensing requirements. Do not reduce this to “all commercial use is forbidden,” and read the current terms for your deployment. SensorFlow's public repository uses Apache-2.0, while its real SDK-ingestion path depends on a separately supplied license.&lt;/p&gt;

&lt;h3&gt;
  
  
  A compact decision matrix
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SDK boundary:&lt;/strong&gt; Countly uses Countly SDKs or its documented write API. SensorFlow is built around compatible existing Sensors Data SDK traffic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data store:&lt;/strong&gt; Countly's server architecture uses MongoDB. SensorFlow's event rows go to your ClickHouse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reporting:&lt;/strong&gt; Countly supplies a product interface. SensorFlow supplies SQL/Superset workflow; your team owns metric definitions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Advanced analysis:&lt;/strong&gt; Countly's current matrix lists built-in funnels and retention for Flex/Enterprise. With SensorFlow, you must implement those definitions in SQL or another analysis layer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Activation:&lt;/strong&gt; Countly Lite has public source with its stated license conditions. SensorFlow's public demo works without a license; real ingestion does not.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Raw data and operational trade-offs
&lt;/h2&gt;

&lt;p&gt;Countly can be self-hosted, and &lt;a href="https://support.countly.com/hc/en-us/articles/900004373266-How-Countly-Works" rel="noopener noreferrer"&gt;its architecture documentation&lt;/a&gt; says stored raw data can be retrieved through REST APIs or MongoDB commands. It would be wrong to imply that only SensorFlow permits raw-data access. The real distinction is whether you prefer Countly's MongoDB model and application workflows or need events to land directly in an existing ClickHouse environment.&lt;/p&gt;

&lt;p&gt;ClickHouse is not automatically faster or cheaper for your workload. SensorFlow puts analytical work on the operator: event schema, identity stitching, deduplication, time zones, retention, backups, access control, and SQL definitions. Countly supplies more application functionality out of the box, but requires its own deployment, SDK, and edition decisions. Test on your own representative event volume and business questions before ranking reliability or total cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  A useful proof of concept
&lt;/h2&gt;

&lt;p&gt;Ask one question: “Of users who started registration, how many completed it within 24 hours?” Agree on event names, stable user identity, time zone, window, and denominator before comparing dashboards.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Countly path:&lt;/strong&gt; Instrument a staging app with the relevant Countly SDK. Send registration-start and registration-complete events. Check that their properties appear in Countly's UI or read API. If you need a built-in funnel, verify your chosen edition includes one; do not assume Lite does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SensorFlow path:&lt;/strong&gt; Launch the demo, inspect sample rows, then activate licensed ingestion for a meaningful test. Send events from your actual Sensors Data SDK version. Query the resulting ClickHouse rows and verify event name, identity, timestamp, and property types. An HTTP success alone does not prove correct storage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recovery path:&lt;/strong&gt; In a test environment, interrupt the store and inspect retries, duplicates, buffering, and restore behavior. A one-event demo is not a production reliability claim.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Analyst path:&lt;/strong&gt; Have the person who will own the metric answer the registration question. Count the work required to build and maintain the answer, not only the time required to start containers.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/docs/getting-started.md" rel="noopener noreferrer"&gt;SensorFlow quick start&lt;/a&gt; documents the demo-first and activation sequence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision
&lt;/h2&gt;

&lt;p&gt;Choose Countly when you want a self-hostable analytics application across web, mobile, or desktop, a built-in reporting workflow, and can adopt Countly instrumentation. Evaluate the Lite versus Flex/Enterprise boundary if funnels, retention, profiles, or engagement features are decisive.&lt;/p&gt;

&lt;p&gt;Evaluate SensorFlow when retaining compatible Sensors Data SDK events and landing them directly in owned ClickHouse are the hard requirements, and a SQL/Superset workflow plus licensed ingestion are acceptable. If you need Countly's integrated UI or a production ingestion path without a separate license, SensorFlow is a poor fit.&lt;/p&gt;

&lt;p&gt;The right question is not “Which has more features?” It is “Which preserves the instrumentation we have, stores data where we need it, and gives our team a reporting workflow it can actually operate?”&lt;/p&gt;

</description>
      <category>analytics</category>
      <category>opensource</category>
      <category>database</category>
    </item>
    <item>
      <title>SensorFlow vs Umami in 2026: Website Analytics or SDK-to-ClickHouse?</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Mon, 28 Sep 2026 08:00:00 +0000</pubDate>
      <link>https://dev.to/sensorflow/sensorflow-vs-umami-in-2026-website-analytics-or-sdk-to-clickhouse-50ki</link>
      <guid>https://dev.to/sensorflow/sensorflow-vs-umami-in-2026-website-analytics-or-sdk-to-clickhouse-50ki</guid>
      <description>&lt;p&gt;Umami and SensorFlow both let a team operate analytics infrastructure, but they solve different starting problems. Umami is a self-hostable web analytics application with its own tracker and reporting interface. SensorFlow is a narrower route for compatible existing Sensors Data SDK events to reach a team-owned ClickHouse database, with SQL and Apache Superset used for analysis.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; Choose Umami first if you want to add website analytics, events, campaigns, and conversion reporting without building a SQL-first workflow. Evaluate SensorFlow if you already have Sensors Data SDK instrumentation and the main requirement is to land those events in your own ClickHouse. Umami is an MIT-licensed open-source project. SensorFlow publishes Apache-2.0 source and a license-free demo, but real SDK ingestion requires a SensorFlow license.&lt;/p&gt;

&lt;p&gt;This comparison uses public documentation as of September 2026. It is not a benchmark, pricing quote, or claim that one product is generally better.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the event source, not the dashboard
&lt;/h2&gt;

&lt;p&gt;Umami's normal website setup is to add a website, install its tracker, and read the resulting reports. Its &lt;a href="https://docs.umami.is/docs" rel="noopener noreferrer"&gt;documentation&lt;/a&gt; covers traffic sources, visitor behavior, conversions, and revenue. It also documents custom events, funnels, and retention; describing Umami as pageviews only would be wrong. The &lt;a href="https://docs.umami.is/docs/track-events" rel="noopener noreferrer"&gt;event-tracking guide&lt;/a&gt; shows HTML data attributes and JavaScript calls.&lt;/p&gt;

&lt;p&gt;For a site with the Umami tracker installed, a named event with properties can be sent like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;umami&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;track&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;signup_completed&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;plan&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;pro&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;step&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That call is appropriate for Umami's instrumentation model. It is not a drop-in replacement for a Sensors Data SDK call. Umami also documents a &lt;a href="https://docs.umami.is/docs/api/sending-stats" rel="noopener noreferrer"&gt;POST /api/send endpoint&lt;/a&gt;, but that endpoint has its own payload shape, website ID, and User-Agent requirement. Pointing an existing Sensors Data SDK server URL at it would not translate the protocol. A migration needs an adapter or client-side changes, plus identity testing.&lt;/p&gt;

&lt;p&gt;SensorFlow's starting point is different: retain compatible SDK instrumentation and change the receiver. The &lt;a href="https://github.com/data-analyze-bi/sensorFlow" rel="noopener noreferrer"&gt;public repository&lt;/a&gt; documents this path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Sensors Data SDK events → SensorFlow Go ingestion → ClickHouse → SQL / Superset
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The README separates a license-free demo from licensed real ingestion. The demo starts supporting services, sample events, and Superset dashboards. To process actual SDK traffic, a team must obtain and activate a license and test its exact SDK versions and payloads. No blanket compatibility claim is justified.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the systems differ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Instrumentation.&lt;/strong&gt; Umami uses its own website tracker, JavaScript calls, or documented API. SensorFlow is evaluated as a receiver for compatible Sensors Data SDK requests. If you have no existing Sensors Data SDK dependency, that SensorFlow advantage may be irrelevant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Storage.&lt;/strong&gt; Umami's current &lt;a href="https://docs.umami.is/docs/install" rel="noopener noreferrer"&gt;installation guide&lt;/a&gt; specifies PostgreSQL and offers Docker Compose, a prebuilt image, or source installation. SensorFlow writes event rows to ClickHouse. This is an architectural distinction, not proof that either database wins on speed or cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Analysis interface.&lt;/strong&gt; Umami offers a built-in analytics interface. Its &lt;a href="https://docs.umami.is/docs/funnel" rel="noopener noreferrer"&gt;funnel documentation&lt;/a&gt; describes ordered URL or event steps and conversion windows, and its &lt;a href="https://docs.umami.is/docs/retention" rel="noopener noreferrer"&gt;retention documentation&lt;/a&gt; describes cohorts of returning visitors. SensorFlow uses ClickHouse SQL and Superset dashboards; it does not offer an equivalent native product-analytics interface. Check feature availability in the Umami version and deployment you actually select.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Terms.&lt;/strong&gt; Umami's repository has a public &lt;a href="https://github.com/umami-software/umami/blob/master/LICENSE" rel="noopener noreferrer"&gt;MIT license&lt;/a&gt;. SensorFlow's repository is Apache-2.0, but its real SDK ingestion path requires a separate license. Treat the cost and support terms as part of the decision, not as a footnote.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Umami is the better first test
&lt;/h2&gt;

&lt;p&gt;For a website team that needs traffic sources, campaigns, custom events, funnels, and retention in one interface, Umami is a more direct workflow. Install its tracker on a staging site, send an event, and ask the actual marketer or analyst to answer a question in the UI. Adding a SQL-first architecture may create work without value if the real need is a usable analytics application.&lt;/p&gt;

&lt;p&gt;Self-hosting does not remove operations. The team still needs HTTPS, access control, PostgreSQL backups, upgrades, tracker delivery, and consent handling appropriate to its own jurisdiction and data. Privacy-focused product design is not a universal legal exemption.&lt;/p&gt;

&lt;h2&gt;
  
  
  When SensorFlow is worth a proof of concept
&lt;/h2&gt;

&lt;p&gt;Suppose a team has many Sensors Data SDK calls and cannot rewrite clients all at once. It wants the event rows in a ClickHouse database it operates and has people who can own SQL metrics. Then SensorFlow's receiver boundary is worth testing: can traffic from the actual SDK versions arrive in ClickHouse with the expected identity, timestamp, and property types?&lt;/p&gt;

&lt;p&gt;That is narrower than replacing Umami. Superset can visualize a funnel or retention query, but the team must define the identities, steps, windows, and denominators, maintain the SQL, and explain the metric to users. Owning raw rows also means owning schema changes, sensitive-property handling, retention, permissions, backups, and incidents.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/docs/getting-started.md" rel="noopener noreferrer"&gt;SensorFlow quick start&lt;/a&gt; lets a developer launch the demo first, then activate licensed ingestion. If the requirement is a fully usable no-license self-hosted collection path, Umami's public terms may fit better. If the requirement is existing Sensors Data SDK traffic in ClickHouse, test SensorFlow's licensed path before deciding whether the migration value justifies it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A useful evaluation for both
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;List current SDKs, event names, properties, and anonymous-to-authenticated identity transitions. Do not assume instrumentation compatibility from a feature list.&lt;/li&gt;
&lt;li&gt;Define what a registration means, the time zone, and the funnel window before opening a dashboard.&lt;/li&gt;
&lt;li&gt;For Umami, install the tracker on a non-production page, send a custom event, and confirm the event and properties in its reporting interface. A JavaScript call alone is not proof of storage.&lt;/li&gt;
&lt;li&gt;For SensorFlow, start the demo and inspect sample ClickHouse rows. If real ingestion is needed, activate the license, send one event from the actual SDK, and query it back. Check event name, distinct ID, timestamp, and property types.&lt;/li&gt;
&lt;li&gt;Let the intended users answer the same business question. Compare the product-manager workflow in Umami with the SQL/Superset workflow the team would maintain in SensorFlow.&lt;/li&gt;
&lt;li&gt;In a test environment, exercise database outage, retries, duplicates, backup restore, and credential rotation. An HTTP success response is not a reconciliation report.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Do not generalize a small test into a throughput or cost ranking. Measure a representative volume and operational workload before making a production decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision
&lt;/h2&gt;

&lt;p&gt;Choose Umami when the main task is adding a self-hostable website analytics application with its own instrumentation and reports. Evaluate SensorFlow when retaining compatible Sensors Data SDK traffic and landing it in owned ClickHouse are the actual constraints, and the team accepts licensed ingestion and a SQL/Superset workflow. They can coexist as separate website and product-event pipelines, but that adds governance work rather than eliminating it.&lt;/p&gt;

&lt;p&gt;SensorFlow is an independent project. It is not affiliated with, endorsed by, or certified by Umami or Sensors Data.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>analytics</category>
      <category>clickhouse</category>
    </item>
    <item>
      <title>I Built an Open-Source, Self-Hosted Alternative to PostHog on ClickHouse</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Sun, 27 Sep 2026 21:54:50 +0000</pubDate>
      <link>https://dev.to/sensorflow/i-built-an-open-source-self-hosted-alternative-to-posthog-on-clickhouse-3bcn</link>
      <guid>https://dev.to/sensorflow/i-built-an-open-source-self-hosted-alternative-to-posthog-on-clickhouse-3bcn</guid>
      <description>&lt;p&gt;Per-event pricing can make analytics costs rise with traffic. I built an alternative: &lt;strong&gt;SensorFlow&lt;/strong&gt; — an open-source, self-hosted event analytics stack from collection to SQL and BI, with data staying on your own servers.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Website: &lt;a href="https://sensorflow.site/" rel="noopener noreferrer"&gt;https://sensorflow.site/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/data-analyze-bi/sensorflow" rel="noopener noreferrer"&gt;https://github.com/data-analyze-bi/sensorflow&lt;/a&gt; (Apache-2.0)&lt;/li&gt;
&lt;li&gt;Live demo: &lt;a href="https://superset.sensorflow.site" rel="noopener noreferrer"&gt;https://superset.sensorflow.site&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The problem with per-event pricing
&lt;/h2&gt;

&lt;p&gt;Managed analytics charges by event volume. The more successful your product, the bigger the bill — growth gets taxed. Self-hosting flips the model: you pay for servers, not for success.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architecture
&lt;/h2&gt;

&lt;p&gt;Deliberately boring:&lt;/p&gt;

&lt;p&gt;SDK → Go collector → ClickHouse → SQL / Apache Superset&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Collect&lt;/strong&gt;: Go ingestion service, wire-compatible with Sensors Data SDKs — existing SDK instrumentation may be reused after checking SDK version, encryption, identity, property types, and timestamps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Store&lt;/strong&gt;: ClickHouse, built for analytical queries (funnels, retention) at scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Analyze&lt;/strong&gt;: Raw SQL when you want control, Superset dashboards when you want clicks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Run ./install.sh for the local demo, then obtain a license and run ./activate.sh to enable real SDK ingestion. No Kubernetes required.&lt;/p&gt;

&lt;h2&gt;
  
  
  SensorFlow vs PostHog
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;SensorFlow&lt;/th&gt;
&lt;th&gt;PostHog&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Deployment&lt;/td&gt;
&lt;td&gt;Self-hosted only&lt;/td&gt;
&lt;td&gt;Cloud or self-hosted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Query model&lt;/td&gt;
&lt;td&gt;SQL-first, ClickHouse-native&lt;/td&gt;
&lt;td&gt;Product UI-first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pricing&lt;/td&gt;
&lt;td&gt;Annual license, $699/yr list price&lt;/td&gt;
&lt;td&gt;Usage-based&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Suite extras&lt;/td&gt;
&lt;td&gt;None (no replay/flags)&lt;/td&gt;
&lt;td&gt;Replay, flags, experiments&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;PostHog is the broader suite. SensorFlow is the narrower, SQL-first option for teams that want ClickHouse and full ownership of their stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing
&lt;/h2&gt;

&lt;p&gt;Annual license starts at &lt;strong&gt;$699/yr&lt;/strong&gt; list price, with the first order currently discounted to &lt;strong&gt;$349&lt;/strong&gt;. It is not billed per event. The first month is free. You bring the servers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Play with the &lt;a href="https://superset.sensorflow.site" rel="noopener noreferrer"&gt;live demo&lt;/a&gt; for five minutes, then run it locally. If it's useful, a star helps: &lt;a href="https://github.com/data-analyze-bi/sensorflow" rel="noopener noreferrer"&gt;https://github.com/data-analyze-bi/sensorflow&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Independent open-source project, not affiliated with Sensors Data.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>clickhouse</category>
      <category>analytics</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>SensorFlow vs Plausible in 2026: Two Very Different Ways to Own Analytics</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Sun, 27 Sep 2026 08:00:00 +0000</pubDate>
      <link>https://dev.to/sensorflow/sensorflow-vs-plausible-in-2026-two-very-different-ways-to-own-analytics-50h4</link>
      <guid>https://dev.to/sensorflow/sensorflow-vs-plausible-in-2026-two-very-different-ways-to-own-analytics-50h4</guid>
      <description>&lt;p&gt;Both Plausible and SensorFlow can put analytics data on infrastructure you operate. Both involve ClickHouse. That is where the easy equivalence ends. Plausible is a privacy-first web analytics product with a tracking script, goals, and a purpose-built dashboard. SensorFlow is a narrower ingestion route for compatible, already-deployed Sensors Data SDK events: SDK → Go receiver → your ClickHouse → Apache Superset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; If your problem is “I need readable website traffic and conversion reports,” start with Plausible. If your problem is “I already have Sensors Data SDK instrumentation and need to test landing its events in my own ClickHouse,” evaluate SensorFlow. Do not choose SensorFlow because someone says Plausible cannot self-host or cannot access raw ClickHouse data: &lt;a href="https://github.com/plausible/analytics" rel="noopener noreferrer"&gt;Plausible's own repository&lt;/a&gt; documents both. And do not assume SensorFlow's open-source demo means production SDK ingestion is free: &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/docs/getting-started.md" rel="noopener noreferrer"&gt;its getting-started guide&lt;/a&gt; separates demo deployment from licensed activation.&lt;/p&gt;

&lt;p&gt;This comparison is based on vendor documentation available in September 2026. It contains no speed, cost, or privacy-compliance benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first distinction is the event contract
&lt;/h2&gt;

&lt;p&gt;Plausible's official &lt;a href="https://plausible.io/docs/events-api" rel="noopener noreferrer"&gt;Events API&lt;/a&gt; accepts POST /api/event with fields including domain, name, and url. Its JavaScript tracker can collect pageviews and custom events, and &lt;a href="https://plausible.io/docs/custom-props/for-custom-events" rel="noopener noreferrer"&gt;custom properties&lt;/a&gt; give those events context. That is a useful model for website teams. The dashboard organizes traffic sources, pages, goals, and conversions without asking an analyst to write SQL.&lt;/p&gt;

&lt;p&gt;SensorFlow's documented starting point is different: existing Sensors Data SDK calls. It attempts to receive compatible SDK payloads through a Go service and write resulting events into ClickHouse. Its core value is preserving a &lt;em&gt;tested subset&lt;/em&gt; of existing instrumentation while changing the server-side destination. That does not make its receiver a drop-in Plausible Events API endpoint. Nor does Plausible's Events API mean it automatically understands a Sensors Data SDK payload. A migration between the two needs an explicit payload mapping, identity decision, and verification of timestamps and properties.&lt;/p&gt;

&lt;p&gt;Consider a practical example: a product already has mobile and web SDK calls for SignupStarted and SignupCompleted, including a stable account identifier. If the team wants the original events for joins and custom SQL, it should test the SensorFlow path with its exact SDK version. If a marketing team chiefly needs privacy-oriented website acquisition and goal reports, it may be faster to instrument Plausible's tracker and use its interface. The event contract, not the word “open source,” decides the implementation work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plausible also uses ClickHouse
&lt;/h2&gt;

&lt;p&gt;It would be misleading to frame this as “Plausible's closed database versus SensorFlow's ClickHouse.” &lt;a href="https://github.com/plausible/analytics" rel="noopener noreferrer"&gt;Plausible's repository&lt;/a&gt; states that its self-hosted Community Edition uses ClickHouse for analytics and PostgreSQL for other application data. It also explicitly says self-hosting allows direct access to raw data in ClickHouse. Its &lt;a href="https://github.com/plausible/community-edition" rel="noopener noreferrer"&gt;Community Edition deployment repository&lt;/a&gt; provides a Docker Compose route.&lt;/p&gt;

&lt;p&gt;The more useful contrast is &lt;em&gt;data model and default workflow&lt;/em&gt;. Plausible's ClickHouse supports Plausible's web analytics product; users normally work in its dashboard and API. SensorFlow's target is a team-operated event store whose rows are queried with ClickHouse SQL and visualized in Superset. If your organization wants to own metric definitions and join events to other warehouse data, that SQL-first approach may suit it. If the organization wants a finished, simple reporting interface, operating ClickHouse alone does not provide one.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;th&gt;Plausible&lt;/th&gt;
&lt;th&gt;SensorFlow&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Primary task&lt;/td&gt;
&lt;td&gt;Website and conversion analytics&lt;/td&gt;
&lt;td&gt;Compatible Sensors Data SDK event ingestion into ClickHouse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud option&lt;/td&gt;
&lt;td&gt;Managed Plausible Cloud&lt;/td&gt;
&lt;td&gt;Self-hosted project; check current commercial license terms for ingestion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-host option&lt;/td&gt;
&lt;td&gt;Plausible Community Edition&lt;/td&gt;
&lt;td&gt;Docker-based demo and licensed real-ingestion path&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analytics store&lt;/td&gt;
&lt;td&gt;ClickHouse in self-hosted CE; PostgreSQL for application data&lt;/td&gt;
&lt;td&gt;Team-operated ClickHouse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Default interface&lt;/td&gt;
&lt;td&gt;Plausible dashboard, goals, reports, APIs&lt;/td&gt;
&lt;td&gt;ClickHouse SQL and Apache Superset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Existing Sensors Data SDK&lt;/td&gt;
&lt;td&gt;Requires integration or mapping to Plausible's event model&lt;/td&gt;
&lt;td&gt;The principal compatibility use case, subject to testing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Built-in funnel UI&lt;/td&gt;
&lt;td&gt;Available in eligible managed plans, not Plausible CE&lt;/td&gt;
&lt;td&gt;No equivalent built-in no-code UI; model it in SQL/Superset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Raw-data access&lt;/td&gt;
&lt;td&gt;Plausible CE allows direct ClickHouse access; Cloud has export/API options&lt;/td&gt;
&lt;td&gt;Direct ClickHouse queries are the normal workflow&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;a href="https://github.com/plausible/analytics" rel="noopener noreferrer"&gt;Plausible README&lt;/a&gt; distinguishes Cloud from CE. In particular, it lists marketing funnels, ecommerce revenue goals, SSO, and Sites API among premium features absent from CE. Its &lt;a href="https://plausible.io/docs/funnel-analysis" rel="noopener noreferrer"&gt;funnel documentation&lt;/a&gt; labels the feature as a Business-plan feature. This edition boundary matters, but it is not a claim that SensorFlow has a better ready-made funnel: it does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Privacy is not a one-word score
&lt;/h2&gt;

&lt;p&gt;Plausible's product is intentionally privacy-first and designed around cookie-free website measurement. Its &lt;a href="https://github.com/plausible/analytics" rel="noopener noreferrer"&gt;repository&lt;/a&gt; describes how it avoids storing personal data and persistent visitor identifiers. The Events API documentation adds an operational nuance: User-Agent and client IP information are used to derive unique visitor counts, and a misconfigured server-side X-Forwarded-For can silently drop an event even if the API returns HTTP 202. A production implementation should follow Plausible's exact guidance rather than equate “202” with a verified stored event.&lt;/p&gt;

&lt;p&gt;SensorFlow can retain event identifiers and properties sent by an existing SDK in a ClickHouse database operated by the customer. That can enable deeper event-level analysis, but it also puts more privacy, consent, access-control, retention, and deletion responsibility on the operator. It would be wrong to market that as inherently more privacy-preserving than Plausible. The two systems are deliberately optimizing different analytics models.&lt;/p&gt;

&lt;h2&gt;
  
  
  The license and operations boundary
&lt;/h2&gt;

&lt;p&gt;Plausible CE is &lt;a href="https://github.com/plausible/analytics" rel="noopener noreferrer"&gt;AGPL-licensed&lt;/a&gt; and self-hostable without a Plausible subscription, while Plausible Cloud is a paid managed service. CE operators still pay for compute, storage, backups, maintenance, and incident response. Cloud includes a managed product and features based on the chosen plan.&lt;/p&gt;

&lt;p&gt;SensorFlow's &lt;a href="https://github.com/data-analyze-bi/sensorFlow" rel="noopener noreferrer"&gt;README&lt;/a&gt; documents an install-first workflow: supporting services and sample dashboards can run before a license is purchased, but receiving real Sensors Data SDK traffic requires a license and activation. This is a material difference from saying “clone the repository and all production ingestion works for free.” Evaluate the license alongside infrastructure and engineering costs.&lt;/p&gt;

&lt;p&gt;Do not confuse &lt;em&gt;running a demo&lt;/em&gt; with a production validation. For SensorFlow, inspect the sample events, then activate an appropriate license and send a real event from the SDK version you actually use. For Plausible, verify the chosen deployment edition, track an event with its documented API or tracker, create the required goal, and confirm that it appears in the dashboard. For both, test proxies, bot filtering, retries, backups, and access to the stored data.&lt;/p&gt;

&lt;h2&gt;
  
  
  A fair proof of concept
&lt;/h2&gt;

&lt;p&gt;Choose one concrete question such as “How many visitors reach signup, and what properties distinguish the people who complete it?” Then:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Write down the source events and identities. Website sessions and authenticated product users are not interchangeable units.&lt;/li&gt;
&lt;li&gt;Track the same &lt;em&gt;business outcome&lt;/em&gt; in both tools. In Plausible, use its documented custom event and goal configuration. In SensorFlow, send a compatible SDK event and query the resulting ClickHouse row.&lt;/li&gt;
&lt;li&gt;Check actual data arrival. Inspect Plausible's dropped-event header where relevant, and query ClickHouse directly for SensorFlow. Do not stop at an HTTP success response.&lt;/li&gt;
&lt;li&gt;Ask the intended users to answer the question. A marketer may prefer Plausible's ready-made dashboard; a data engineer may prefer raw rows and SQL. Neither preference is universal.&lt;/li&gt;
&lt;li&gt;Record setup and maintenance work, product entitlements, and the cost of changing existing SDK code. Do not compare a managed Plausible Cloud plan against an unoperated SensorFlow demo as though they provide the same service.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Which should you choose?
&lt;/h2&gt;

&lt;p&gt;Use Plausible when you want privacy-first website analytics with a focused dashboard and tracker. Its Community Edition is a real self-hosted option with ClickHouse and raw-data access, although the CE feature set differs from paid Cloud plans.&lt;/p&gt;

&lt;p&gt;Test SensorFlow when you already have compatible Sensors Data SDK instrumentation, explicitly want that event stream in your own ClickHouse, and have the engineering capacity to define and operate the analyses. Its narrower scope and licensed real-ingestion boundary should be stated upfront. If you need a mature out-of-the-box website analytics application, Plausible is usually the better starting point.&lt;/p&gt;

&lt;p&gt;SensorFlow is an independent project. It is not affiliated with, endorsed by, or certified by Plausible or Sensors Data.&lt;/p&gt;

</description>
      <category>analytics</category>
      <category>opensource</category>
      <category>database</category>
      <category>architecture</category>
    </item>
    <item>
      <title>SensorFlow vs Matomo in 2026: Event Pipeline or Web Analytics?</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Sat, 26 Sep 2026 12:50:05 +0000</pubDate>
      <link>https://dev.to/sensorflow/sensorflow-vs-matomo-in-2026-event-pipeline-or-web-analytics-2f7p</link>
      <guid>https://dev.to/sensorflow/sensorflow-vs-matomo-in-2026-event-pipeline-or-web-analytics-2f7p</guid>
      <description>&lt;p&gt;Matomo and SensorFlow are both relevant when a team wants to run analytics on its own infrastructure. They are not substitutes for the same job. Matomo is a mature web analytics application with tracking, reports, dashboards, and a self-hosted edition. SensorFlow is a narrower path for teams that already send compatible Sensors Data SDK events and want those events in their own ClickHouse, queried through SQL and Apache Superset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; Pick Matomo if you need a ready-to-use website analytics interface, visitor reports, campaigns, goals, and event tracking. Evaluate SensorFlow if your starting point is existing Sensors Data SDK instrumentation and your desired destination is ClickHouse. SensorFlow's demo stack can be run before purchase, but activating real SDK ingestion requires a SensorFlow license. Neither tool should be presented as a free, feature-equivalent replacement for the other.&lt;/p&gt;

&lt;p&gt;This comparison uses public documentation as of September 2026. It is not a performance benchmark or a claim that one product is universally better.&lt;/p&gt;

&lt;h2&gt;
  
  
  What are you actually replacing?
&lt;/h2&gt;

&lt;p&gt;The useful question is not merely whether both projects can be self-hosted. Ask where the boundary you need to change sits.&lt;/p&gt;

&lt;p&gt;If your team currently looks at website acquisition, visits, content performance, campaigns, and conversion goals, the boundary is probably the analytics application. Matomo already provides that application, plus an official event-tracking workflow. Its &lt;a href="https://matomo.org/guide/installation-maintenance/matomo-on-premise-self-hosted/" rel="noopener noreferrer"&gt;on-premise installation guide&lt;/a&gt; says the self-hosted edition is free to download. The &lt;a href="https://matomo.org/guide/reports/event-tracking/" rel="noopener noreferrer"&gt;event-tracking guide&lt;/a&gt; describes the interaction data it can collect.&lt;/p&gt;

&lt;p&gt;If your team already has Sensors Data SDK calls across web, mobile, or backend code, the boundary may instead be the event receiver. SensorFlow's documented flow is compatible SDK traffic → Go ingestion → ClickHouse → Superset. The objective is to test a server-side destination change without rewriting all instrumentation at once. See the &lt;a href="https://github.com/data-analyze-bi/sensorFlow" rel="noopener noreferrer"&gt;SensorFlow repository&lt;/a&gt; and &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/docs/getting-started.md" rel="noopener noreferrer"&gt;getting-started guide&lt;/a&gt; for the actual deployment and license steps.&lt;/p&gt;

&lt;p&gt;That narrower scope matters. SensorFlow does not turn Superset into Matomo's visitor interface, and Matomo does not natively make an existing Sensors Data SDK send its events to ClickHouse. A team can build additional integrations, but that work belongs in the comparison.&lt;/p&gt;

&lt;h2&gt;
  
  
  Side-by-side: where each tool fits
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;th&gt;Matomo&lt;/th&gt;
&lt;th&gt;SensorFlow&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Primary job&lt;/td&gt;
&lt;td&gt;Web analytics application with built-in reports and tracking&lt;/td&gt;
&lt;td&gt;Server-side path from compatible Sensors Data SDK events to ClickHouse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-hosting&lt;/td&gt;
&lt;td&gt;Official On-Premise edition&lt;/td&gt;
&lt;td&gt;Docker-based stack and self-operated ingestion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical storage&lt;/td&gt;
&lt;td&gt;MySQL or MariaDB for Matomo On-Premise&lt;/td&gt;
&lt;td&gt;ClickHouse for event rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analysis interface&lt;/td&gt;
&lt;td&gt;Matomo reports and dashboards&lt;/td&gt;
&lt;td&gt;ClickHouse SQL and Apache Superset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Raw-data access&lt;/td&gt;
&lt;td&gt;API export; on-premise users can also read the underlying database&lt;/td&gt;
&lt;td&gt;Direct ClickHouse queries are the normal workflow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Existing Sensors Data SDK&lt;/td&gt;
&lt;td&gt;A separate migration/integration exercise&lt;/td&gt;
&lt;td&gt;The principal compatibility use case, subject to testing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Funnels&lt;/td&gt;
&lt;td&gt;Check plan or on-premise plugin entitlement&lt;/td&gt;
&lt;td&gt;Model in SQL and present in Superset; not a comparable no-code UI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Real ingestion terms&lt;/td&gt;
&lt;td&gt;Depend on hosting, plan, and plugins&lt;/td&gt;
&lt;td&gt;Demo before activation; real SDK ingestion requires a license&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Matomo's &lt;a href="https://matomo.org/faq/on-premise/matomo-requirements/" rel="noopener noreferrer"&gt;system requirements&lt;/a&gt; name PHP and MySQL/MariaDB for its on-premise application. This is an architectural difference, not evidence that one database is faster. Capacity, operational cost, and query performance must be measured on your workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  Matomo is not just page views
&lt;/h2&gt;

&lt;p&gt;It would be inaccurate to dismiss Matomo as a simple traffic counter. Its &lt;a href="https://matomo.org/pricing/" rel="noopener noreferrer"&gt;current feature page&lt;/a&gt; includes event tracking, ecommerce tracking, segmentation, dashboards, API access, and reporting features. Matomo's event-tracking documentation covers interactions such as clicks, videos, downloads, and forms. For many website teams, those built-in workflows are the reason to choose it: an analyst can navigate a product interface without first maintaining a catalog of SQL queries.&lt;/p&gt;

&lt;p&gt;It would also be inaccurate to say that only SensorFlow provides access to raw data. Matomo &lt;a href="https://matomo.org/faq/how-to/faq_24536/" rel="noopener noreferrer"&gt;documents both HTTP API export and direct read-only access to its MySQL database for on-premise installations&lt;/a&gt;. If your requirement is simply to extract your own events, Matomo On-Premise deserves a fair evaluation. The sharper distinction is where raw rows live by default and how the team expects to work with them. SensorFlow is designed around ClickHouse as the event store; Matomo's application and schema are designed around Matomo's reports.&lt;/p&gt;

&lt;p&gt;Some advanced Matomo features depend on a plan or premium plugin. The current on-premise pricing page places Funnels in a paid bundle rather than the free Community feature set. Verify the current bundle, support terms, and deployment model on &lt;a href="https://matomo.org/pricing/" rel="noopener noreferrer"&gt;Matomo's pricing page&lt;/a&gt; before making a purchase decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  SensorFlow is not a complete Matomo replacement
&lt;/h2&gt;

&lt;p&gt;The SensorFlow &lt;a href="https://github.com/data-analyze-bi/sensorFlow" rel="noopener noreferrer"&gt;README&lt;/a&gt; makes a two-stage distinction: the installer starts supporting services, demo events, and a Superset dashboard without a license; real Sensors Data SDK ingestion starts only after a license is installed and activation is run. Calling the entire production ingestion path free would hide a material requirement.&lt;/p&gt;

&lt;p&gt;After activation, the attraction for a SQL-oriented team is that the event receiver writes to a ClickHouse database the team operates. Engineers can inspect event rows, verify property types, define metrics, and join with other data using their own SQL. Superset can present dashboards, but it does not automatically supply Matomo's reporting vocabulary or workflow. The team owns metric definitions, dashboard quality, retention, backup, security, and upgrades.&lt;/p&gt;

&lt;p&gt;The compatibility claim must stay narrow. Working with compatible Sensors Data SDK traffic does not establish parity for every SDK generation, encrypted payload, visual tracking mode, identity rule, or plugin. A proof of concept should use the exact SDK versions and event payloads already deployed.&lt;/p&gt;

&lt;h2&gt;
  
  
  A concrete evaluation
&lt;/h2&gt;

&lt;p&gt;Start with one business question: “Of the people who clicked sign-up this week, how many finished registration?” Run the same decision workflow in each candidate system.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Inventory current instrumentation.&lt;/strong&gt; Record event names, properties, user identifiers, SDK versions, and destinations. If the application uses Matomo's tracker rather than Sensors Data SDKs, SensorFlow's compatibility advantage may be irrelevant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Define the event contract.&lt;/strong&gt; Specify whether “sign-up” means a click, submitted form, created account, or verified account. Neither product can fix an ambiguous definition.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deploy in non-production.&lt;/strong&gt; Follow Matomo's &lt;a href="https://matomo.org/guide/installation-maintenance/matomo-on-premise-self-hosted/" rel="noopener noreferrer"&gt;installation guide&lt;/a&gt; and SensorFlow's &lt;a href="https://github.com/data-analyze-bi/sensorFlow/blob/main/docs/getting-started.md" rel="noopener noreferrer"&gt;getting-started guide&lt;/a&gt;. Keep test tokens, database credentials, and network ports private.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Send identifiable test events.&lt;/strong&gt; Include a unique test user and non-sensitive properties. Confirm both the HTTP response and the stored record or report; an accepted request alone does not prove retention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reconcile counts and identities.&lt;/strong&gt; Compare event counts by day, environment, and event name. Check anonymous-to-authenticated transitions and time zones. Treat divergence as a finding, not a rounding error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Let intended users answer the question.&lt;/strong&gt; Ask a marketer or product manager to find the answer in Matomo, and ask the intended SQL/Superset owner to produce and explain the query. The working experience matters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test failure and recovery.&lt;/strong&gt; Stop the database in a test environment, rotate a token, restore from backup, and see whether lost or delayed events can be diagnosed before production use.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A small demo verifies a path; it does not predict production throughput or total cost. Measure representative volume, query mix, retention, and operations effort before selecting a winner.&lt;/p&gt;

&lt;h2&gt;
  
  
  When should you choose each?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Matomo is the more direct choice&lt;/strong&gt; for a website team that wants a complete analytics interface now. Reports, goals, campaigns, and event tracking already live in the same application. Its on-premise option allows self-hosting, and the free Community tier may cover the required core features. Teams can later evaluate paid plugins and support.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SensorFlow is worth a proof of concept&lt;/strong&gt; when compatible Sensors Data SDK events already exist, ClickHouse is the desired event store, and the team can operate SQL and infrastructure. The first milestone is to start the demo, inspect sample events, activate licensed ingestion if the real path is needed, then send one event from the actual SDK and query it back. Expand gradually only after identity and property types check out.&lt;/p&gt;

&lt;p&gt;Matomo should not be characterized as lacking events or raw-data access. SensorFlow should not be characterized as a complete or freely activated Matomo replacement. Choose based on the boundary you need to change and on who will operate and use the system.&lt;/p&gt;

&lt;p&gt;SensorFlow is an independent project. It is not affiliated with, endorsed by, or certified by Matomo or Sensors Data.&lt;/p&gt;

</description>
      <category>analytics</category>
      <category>opensource</category>
      <category>database</category>
      <category>architecture</category>
    </item>
    <item>
      <title>SensorFlow vs GrowingIO in 2026: Analytics Product or Self-Hosted Event Pipeline?</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Fri, 25 Sep 2026 03:27:53 +0000</pubDate>
      <link>https://dev.to/sensorflow/sensorflow-vs-growingio-in-2026-analytics-product-or-self-hosted-event-pipeline-4k8m</link>
      <guid>https://dev.to/sensorflow/sensorflow-vs-growingio-in-2026-analytics-product-or-self-hosted-event-pipeline-4k8m</guid>
      <description>&lt;p&gt;GrowingIO and SensorFlow address different parts of event analytics. GrowingIO offers collection SDKs and product-analysis workflows, including documented automatic collection and funnel analysis. SensorFlow is a narrower self-hosted pipeline for compatible Sensors Data SDK events: Go collector → your ClickHouse → SQL and Apache Superset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; Evaluate GrowingIO if product and growth teams need ready-made analysis workflows and automatic collection. Evaluate SensorFlow if you already use a Sensors Data SDK, need raw events in your own ClickHouse, and have the engineering capacity to operate the stack. An existing GrowingIO SDK integration is &lt;strong&gt;not&lt;/strong&gt; a drop-in input for SensorFlow.&lt;/p&gt;

&lt;h2&gt;
  
  
  The important differences
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;SensorFlow&lt;/th&gt;
&lt;th&gt;GrowingIO&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Main scope&lt;/td&gt;
&lt;td&gt;Self-hosted collection and event storage&lt;/td&gt;
&lt;td&gt;Collection and product-analysis workflows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Client migration path&lt;/td&gt;
&lt;td&gt;Verified standard Sensors Data SDK event requests&lt;/td&gt;
&lt;td&gt;GrowingIO SDKs and platform configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Automatic collection&lt;/td&gt;
&lt;td&gt;Do not assume feature parity&lt;/td&gt;
&lt;td&gt;Documented for supported SDK versions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Funnels&lt;/td&gt;
&lt;td&gt;Define semantics and write ClickHouse SQL, then visualize in Superset&lt;/td&gt;
&lt;td&gt;Documented funnel-analysis workflow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Raw ClickHouse rows&lt;/td&gt;
&lt;td&gt;Query the team's own database&lt;/td&gt;
&lt;td&gt;Confirm access and export terms for the specific plan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operations&lt;/td&gt;
&lt;td&gt;The deploying team&lt;/td&gt;
&lt;td&gt;Depends on the purchased and deployment arrangement&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;GrowingIO's &lt;a href="https://docs.growingio.com/v3/developer-manual/sdkintegrated/android-sdk/auto-android-sdk" rel="noopener noreferrer"&gt;automatic collection documentation&lt;/a&gt; describes supported page and element events. Its &lt;a href="https://docs.growingio.com/op/product-manual/product-analysis/funnel" rel="noopener noreferrer"&gt;funnel documentation&lt;/a&gt; describes conversion windows and breakdowns. Those are product capabilities SensorFlow should not claim to match with an out-of-the-box interface. GrowingIO also documents an &lt;a href="https://docs.growingio.com/op/developer-manual/api-reference/cdp/guide" rel="noopener noreferrer"&gt;analysis-result export API&lt;/a&gt;, so a blanket claim that its data cannot be exported would be inaccurate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a GrowingIO SDK cannot simply point at SensorFlow
&lt;/h2&gt;

&lt;p&gt;SensorFlow's intended ingestion path is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Verified Sensors Data SDK event request
                 ↓
           Go collector
                 ↓
       Your ClickHouse
                 ↓
         SQL / Superset
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Changing an SDK's server URL only works if the destination understands that SDK's request format and semantics. SensorFlow does not document support for the GrowingIO SDK protocol. Request framing, identity fields, retries, and automatically collected event meanings need a separate migration design. If your analytics relies heavily on GrowingIO's automatic or visual event collection, plan to rebuild and verify those events; they do not appear automatically when you replace the receiving server.&lt;/p&gt;

&lt;p&gt;The migration boundary is different for an app already using a Sensors Data SDK. Test the actual SDK version and payloads against SensorFlow in a non-production environment, verify collector logs, and query the resulting ClickHouse rows before shifting traffic. Extensions, identity merges, and unusual event types still require explicit tests. The &lt;a href="https://github.com/data-analyze-bi/sensorFlow" rel="noopener noreferrer"&gt;SensorFlow repository&lt;/a&gt; contains deployment and verification materials. Demo dashboards can start before activation; receiving real events requires a license.&lt;/p&gt;

&lt;h2&gt;
  
  
  What owning the event table buys you
&lt;/h2&gt;

&lt;p&gt;With events in your ClickHouse instance, you can inspect individual rows instead of relying only on an aggregate chart. After sending a uniquely named test event to an activated collector:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="nb"&gt;time&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;distinct_id&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;sensors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'integration_test'&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="nb"&gt;time&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
&lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then compare daily event and user counts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;toDate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;time&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;event_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;uniqExact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;distinct_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;sensors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="nb"&gt;time&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;INTERVAL&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt; &lt;span class="k"&gt;DAY&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;event_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;event_date&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These queries validate the SensorFlow table, not equivalence with a GrowingIO report. Funnels, retention, identity stitching, and “active user” definitions remain your team's responsibility. So do ClickHouse capacity, backups, access control, HTTPS, monitoring, and upgrades.&lt;/p&gt;

&lt;h2&gt;
  
  
  A useful proof of concept
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Collection coverage:&lt;/strong&gt; Inventory five business-critical events, including any auto-collected or visually configured events. Confirm each candidate produces the same business meaning, not merely a similar count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metric reproduction:&lt;/strong&gt; Ask the people who use the analysis to build a signup funnel, next-day retention, and channel conversion. Record discrepancies and the amount of SQL support required.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operations and rollback:&lt;/strong&gt; Test collector or ClickHouse downtime, payload rejection, and license-processing errors. Define rollout and rollback conditions before sending production traffic.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;GrowingIO is the better fit when a ready-to-use product-analysis interface and collection workflows matter more than owning a particular event-storage path. SensorFlow is worth testing when the team already has SQL, Docker, and ClickHouse expertise—particularly if Sensors Data SDK instrumentation is in place and control of raw events is the primary requirement. For a GrowingIO-SDK-heavy application, SensorFlow is &lt;strong&gt;not&lt;/strong&gt; a low-effort replacement.&lt;/p&gt;

&lt;p&gt;SensorFlow is an independent project and is not affiliated with or endorsed by GrowingIO or Sensors Data. Features, deployment terms, and data-access rights for any commercial plan should be verified against current vendor documentation and the actual agreement.&lt;/p&gt;

</description>
      <category>clickhouse</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>SensorFlow vs PostHog in 2026: Which Fits Your Stack?</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Thu, 24 Sep 2026 14:20:14 +0000</pubDate>
      <link>https://dev.to/sensorflow/sensorflow-vs-posthog-in-2026-which-fits-your-stack-4p83</link>
      <guid>https://dev.to/sensorflow/sensorflow-vs-posthog-in-2026-which-fits-your-stack-4p83</guid>
      <description>&lt;h1&gt;
  
  
  SensorFlow vs PostHog in 2026: Which Fits Your Stack?
&lt;/h1&gt;

&lt;p&gt;SensorFlow and PostHog can both appear in a search for self-hosted event analytics, but they are not equivalent products. PostHog is an integrated product-engineering platform. SensorFlow is a narrower event-data path for teams that want compatible Sensors Data SDK traffic to flow through a Go collector into their own ClickHouse and Apache Superset.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Quick answer:&lt;/strong&gt; Choose PostHog when product analytics, session replay, feature flags, experiments, and a unified product interface matter more than assembling your own data workflow. Evaluate SensorFlow when retaining existing Sensors Data SDK instrumentation, owning raw ClickHouse events, and defining metrics in SQL are the primary requirements.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The practical difference
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decision area&lt;/th&gt;
&lt;th&gt;SensorFlow&lt;/th&gt;
&lt;th&gt;PostHog&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Primary scope&lt;/td&gt;
&lt;td&gt;Self-hosted event ingestion and analytics data path&lt;/td&gt;
&lt;td&gt;Integrated product analytics and product-engineering suite&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical flow&lt;/td&gt;
&lt;td&gt;Sensors Data SDK → Go → ClickHouse → Superset&lt;/td&gt;
&lt;td&gt;PostHog SDKs/events → PostHog analytics surfaces&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Product analytics UI&lt;/td&gt;
&lt;td&gt;Superset dashboards and SQL workflows&lt;/td&gt;
&lt;td&gt;Native trends, funnels, retention, paths, stickiness, and lifecycle insights&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Session replay&lt;/td&gt;
&lt;td&gt;Not a core built-in capability&lt;/td&gt;
&lt;td&gt;Officially documented product capability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feature flags and experiments&lt;/td&gt;
&lt;td&gt;Not a core built-in capability&lt;/td&gt;
&lt;td&gt;Integrated with the product suite&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Raw-data workflow&lt;/td&gt;
&lt;td&gt;Direct ClickHouse access is central&lt;/td&gt;
&lt;td&gt;Product UI and supported query surfaces are central&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operational model&lt;/td&gt;
&lt;td&gt;Team operates the components&lt;/td&gt;
&lt;td&gt;Cloud and self-hosting options must be evaluated against current official guidance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best fit&lt;/td&gt;
&lt;td&gt;Data and engineering teams with existing compatible SDK traffic&lt;/td&gt;
&lt;td&gt;Product, growth, and engineering teams seeking an integrated workflow&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;PostHog's current product analytics documentation describes trends, funnels, retention, paths, stickiness, and lifecycle insights. It also connects those events to session replay, feature flags, and experiments. That breadth is a real advantage for teams that want to move from observing a conversion change to inspecting sessions or controlling a rollout in one product.&lt;/p&gt;

&lt;p&gt;Official reference: &lt;a href="https://posthog.com/docs/product-analytics" rel="noopener noreferrer"&gt;https://posthog.com/docs/product-analytics&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where PostHog is the stronger choice
&lt;/h2&gt;

&lt;p&gt;PostHog is usually the stronger candidate when product managers and growth teams need to explore behavior without waiting for a new SQL query. Native analytics concepts reduce the amount of dashboard modeling required before a team can inspect funnels or retention.&lt;/p&gt;

&lt;p&gt;It is also the clearer choice when replay, flags, and experiments are requirements rather than optional adjacent tools. SensorFlow should not claim parity in those areas. Reconstructing them with SQL and a BI layer would not provide the same workflow or product experience.&lt;/p&gt;

&lt;p&gt;Teams should still read PostHog's current deployment and licensing documentation instead of relying on old comparison posts. “Source available,” “open source,” “self-hostable,” and “all features under one license” are different claims, and deployment recommendations can change over time.&lt;/p&gt;

&lt;p&gt;Self-hosting reference: &lt;a href="https://posthog.com/docs/self-host" rel="noopener noreferrer"&gt;https://posthog.com/docs/self-host&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where SensorFlow is meaningfully different
&lt;/h2&gt;

&lt;p&gt;SensorFlow starts with a specific migration problem: a team already sends events through official Sensors Data SDKs and wants to change the server-side destination before rewriting every client integration.&lt;/p&gt;

&lt;p&gt;Its public architecture is intentionally simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Compatible Sensors Data SDK traffic
              ↓
          Go collector
              ↓
          ClickHouse
              ↓
     SQL + Apache Superset
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For this audience, direct ownership of the ClickHouse rows is not an export feature; it is the normal operating model. Engineers can inspect event schemas, query raw records, join events with internal data, and decide how dashboards are calculated.&lt;/p&gt;

&lt;p&gt;That flexibility comes with responsibility. The team must operate HTTPS, authentication, ClickHouse storage, backups, monitoring, upgrades, and metric definitions. SensorFlow is not the easier choice for every organization merely because its component list is smaller.&lt;/p&gt;

&lt;p&gt;Project source: &lt;a href="https://github.com/data-analyze-bi/sensorFlow" rel="noopener noreferrer"&gt;https://github.com/data-analyze-bi/sensorFlow&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Is SensorFlow a PostHog replacement?
&lt;/h2&gt;

&lt;p&gt;Not in the general sense. It can replace a narrower portion of an event pipeline for teams whose main objective is self-hosted collection, ClickHouse storage, and SQL-driven analysis. It does not replace PostHog's complete product analytics interface, session replay, feature-flag, or experimentation workflows.&lt;/p&gt;

&lt;p&gt;The phrase “PostHog alternative” is only useful after the required scope is stated. If the scope is “an integrated product stack,” SensorFlow is not feature-equivalent. If the scope is “a transparent event path into our own ClickHouse,” SensorFlow may be a relevant alternative to evaluate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Migration and validation considerations
&lt;/h2&gt;

&lt;p&gt;Whichever system a team selects, a production decision should use real events rather than a feature checklist. A useful proof of concept should verify:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;SDK request compatibility and retry behavior.&lt;/li&gt;
&lt;li&gt;Anonymous and authenticated identity handling.&lt;/li&gt;
&lt;li&gt;Timestamp, numeric, boolean, array, and nested-property semantics.&lt;/li&gt;
&lt;li&gt;Event-count reconciliation between old and new destinations.&lt;/li&gt;
&lt;li&gt;Query latency on representative retention and funnel workloads.&lt;/li&gt;
&lt;li&gt;Backup, restore, monitoring, and upgrade procedures.&lt;/li&gt;
&lt;li&gt;Access control and handling of sensitive properties.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;SensorFlow's compatibility claims should be tested against the exact Sensors Data SDK versions and extensions in use. Encryption plugins, visual tracking, autocapture variants, and identity behavior should not be assumed compatible without evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Total cost is more than hosting
&lt;/h2&gt;

&lt;p&gt;A self-hosted stack has infrastructure costs, but engineering time is often the more important variable. SensorFlow requires people who can operate Docker, ClickHouse, SQL, and a BI layer. PostHog's broader stack can also require meaningful operational work when self-hosted, while its hosted offering changes that responsibility.&lt;/p&gt;

&lt;p&gt;Compare the full operating model: software terms, compute and storage, backups, upgrades, incident response, analyst time, product-manager autonomy, and the cost of maintaining metric definitions. Avoid unsourced performance or price claims; measure the workload that resembles your own product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who should choose which?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Choose PostHog first&lt;/strong&gt; when an integrated product analytics experience is the goal; product teams need native funnels and retention; session replay is required; or feature flags and experiments should share the same events and identities.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evaluate SensorFlow first&lt;/strong&gt; when a team already uses compatible Sensors Data SDK traffic; raw ClickHouse ownership is mandatory; SQL and Superset are accepted working tools; and the migration should begin at the server-side ingestion boundary.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Choose neither without a proof of concept&lt;/strong&gt; when identity rules are complex, regulatory constraints are strict, event volume is large, or critical SDK extensions are involved. Those conditions require measured compatibility and operational testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;PostHog is the more complete product-engineering platform. SensorFlow is the more focused data-path option. The correct choice depends on whether the team wants an integrated product workflow or a transparent, self-operated route from existing compatible SDK events into ClickHouse.&lt;/p&gt;

&lt;p&gt;SensorFlow is an independent project and is not affiliated with or endorsed by PostHog or Sensors Data.&lt;/p&gt;

</description>
      <category>selfhosted</category>
      <category>clickhouse</category>
    </item>
    <item>
      <title>How to Validate Event Data in ClickHouse and Superset</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Fri, 18 Sep 2026 18:05:51 +0000</pubDate>
      <link>https://dev.to/sensorflow/how-to-validate-event-data-in-clickhouse-and-superset-4369</link>
      <guid>https://dev.to/sensorflow/how-to-validate-event-data-in-clickhouse-and-superset-4369</guid>
      <description>&lt;p&gt;An analytics migration is not complete when an ingestion endpoint returns HTTP 200. That response proves only that one request reached one service. It does not prove that the event was decoded correctly, stored with the intended types, associated with the right identity, interpreted in the right time zone, or counted consistently in a dashboard.&lt;/p&gt;

&lt;p&gt;The safer approach is to validate the pipeline in layers: transport, raw storage, semantic fields, aggregate queries, and dashboard output. This article presents a repeatable checklist for teams using ClickHouse and Apache Superset. SensorFlow is used as an Apache-2.0 implementation example, but the method applies to other self-hosted event pipelines.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Direct answer:&lt;/strong&gt; Validate event data at three minimum checkpoints: confirm the ingestion response and service logs, query the exact raw event in ClickHouse, then reproduce the intended metric in Superset. A successful request is necessary, but only agreement across all three layers shows that the event is usable for analysis.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why a 200 response is not enough
&lt;/h2&gt;

&lt;p&gt;An event can pass through the network while still becoming analytically wrong. A decoder may accept the payload but drop an unsupported property. A timestamp can be interpreted in the wrong time zone. A numeric value can arrive as a string. A retry can create a duplicate. An identity transition can split one person into two users.&lt;/p&gt;

&lt;p&gt;These errors are especially dangerous because dashboards may still look plausible. The total can be close enough that nobody notices until a business decision depends on a segment, funnel, or retention calculation.&lt;/p&gt;

&lt;p&gt;ClickHouse describes itself as a column-oriented SQL database for online analytical processing. Its &lt;a href="https://clickhouse.com/docs" rel="noopener noreferrer"&gt;official documentation&lt;/a&gt; explains the database behavior that should guide schema and query design. Apache Superset is a separate exploration and visualization layer; its &lt;a href="https://superset.apache.org/docs/intro/" rel="noopener noreferrer"&gt;official introduction&lt;/a&gt; covers SQL exploration, datasets, charts, and dashboards. Neither component can infer your event semantics automatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five-layer validation model
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Transport&lt;/td&gt;
&lt;td&gt;Did the endpoint receive and accept the request?&lt;/td&gt;
&lt;td&gt;HTTP response, request ID, ingestion log&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;Is the exact event present in ClickHouse?&lt;/td&gt;
&lt;td&gt;Raw-row query using a unique test marker&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantics&lt;/td&gt;
&lt;td&gt;Are identity, time, name, and properties correct?&lt;/td&gt;
&lt;td&gt;Field-by-field comparison with the sent payload&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aggregation&lt;/td&gt;
&lt;td&gt;Does SQL produce the expected count and distinct users?&lt;/td&gt;
&lt;td&gt;Saved validation query with a narrow time range&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Presentation&lt;/td&gt;
&lt;td&gt;Does Superset show the same result?&lt;/td&gt;
&lt;td&gt;Dataset SQL, filters, time zone, and chart output&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each layer should be checked independently. If a chart is wrong, starting with the chart often wastes time; first locate the raw event, then move upward through the query and dataset configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Create an identifiable test event
&lt;/h2&gt;

&lt;p&gt;Do not begin with anonymous production traffic. Send one event that can be found without ambiguity. Use a name such as &lt;code&gt;integration_test&lt;/code&gt; and include a unique run identifier, test environment, client platform, SDK version, and a known test user.&lt;/p&gt;

&lt;p&gt;The test should contain representative property types:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a string such as &lt;code&gt;plan_name&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;an integer such as &lt;code&gt;item_count&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;a decimal if the protocol supports one;&lt;/li&gt;
&lt;li&gt;a Boolean such as &lt;code&gt;is_trial&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;a timestamp with an explicit offset;&lt;/li&gt;
&lt;li&gt;an array or nested value only if the receiving protocol claims to support it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Record the payload before sending it. That payload becomes the expected result for the semantic comparison. Never include real personal data in a migration test when synthetic identifiers are sufficient.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Verify transport and decoding
&lt;/h2&gt;

&lt;p&gt;Capture the HTTP status, response body, request time, destination URL, and any request identifier returned by the service. Then inspect the ingestion log for the same run identifier.&lt;/p&gt;

&lt;p&gt;A useful transport check answers four separate questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Did the request reach the intended environment?&lt;/li&gt;
&lt;li&gt;Did authentication or signature validation pass?&lt;/li&gt;
&lt;li&gt;Did the decoder recognize the event format?&lt;/li&gt;
&lt;li&gt;Did the storage operation complete rather than enter a retry or dead-letter path?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Treat parsing warnings as failures until reviewed. Silently discarded properties are data loss even when the top-level event is stored.&lt;/p&gt;

&lt;p&gt;SensorFlow focuses on this receiving boundary: a Go service accepts a compatible standard event upload flow and writes events to ClickHouse for SQL inspection. It is an independent project and is not affiliated with or endorsed by Sensors Data. Compatibility must be verified for the SDK versions, plugins, encryption settings, and upload modes used by a specific deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Find the raw event in ClickHouse
&lt;/h2&gt;

&lt;p&gt;Query the narrowest possible time range and filter by the unique run identifier. Adapt database, table, and column names to the deployed schema:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;distinct_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;event_time&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;received_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;properties&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;sensors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'integration_test'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;properties&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'validation_run'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'2026-09-17-a'&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;received_at&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
&lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The expected result is usually one row. Zero rows indicate a transport-to-storage problem. Multiple rows may be legitimate retries, but they require an explicit deduplication policy rather than an assumption.&lt;/p&gt;

&lt;p&gt;For production schemas based on the MergeTree family, study ClickHouse's &lt;a href="https://clickhouse.com/docs/engines/table-engines/mergetree-family/mergetree" rel="noopener noreferrer"&gt;MergeTree documentation&lt;/a&gt;. The sorting key, partition strategy, and common filters determine how much data a query reads. A validation query that works on a tiny test table can become expensive if the production table design does not match the access pattern.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Quotable rule:&lt;/strong&gt; Keep the raw event layer queryable. When a metric is disputed, the team should be able to move from a dashboard number to its SQL definition and then to the underlying events without relying on a closed transformation.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Step 4: Compare semantic fields
&lt;/h2&gt;

&lt;p&gt;Finding the row is only the beginning. Compare the stored record with the original payload field by field.&lt;/p&gt;

&lt;h3&gt;
  
  
  Event name
&lt;/h3&gt;

&lt;p&gt;Check exact spelling and case. Decide whether the system normalizes names and document that behavior. &lt;code&gt;SignUp&lt;/code&gt;, &lt;code&gt;signup&lt;/code&gt;, and &lt;code&gt;sign_up&lt;/code&gt; should not accidentally become three business events.&lt;/p&gt;

&lt;h3&gt;
  
  
  Identity
&lt;/h3&gt;

&lt;p&gt;Inspect anonymous ID, login ID, device ID, and any identity association event. A successful login transition should not unexpectedly split pre-login and post-login behavior. Define whether counts use events, devices, anonymous IDs, account IDs, or a resolved person ID.&lt;/p&gt;

&lt;h3&gt;
  
  
  Time
&lt;/h3&gt;

&lt;p&gt;Store enough information to distinguish client occurrence time from server receipt time. Verify time zone conversion and test around midnight, daylight-saving transitions where relevant, delayed uploads, and offline queues.&lt;/p&gt;

&lt;h3&gt;
  
  
  Property types
&lt;/h3&gt;

&lt;p&gt;Confirm that numeric, Boolean, date, string, and collection values preserve their intended types. A property that changes from &lt;code&gt;42&lt;/code&gt; to &lt;code&gt;"42"&lt;/code&gt; can break filters, sorting, ranges, and aggregations without making the event disappear.&lt;/p&gt;

&lt;h3&gt;
  
  
  Environment and source
&lt;/h3&gt;

&lt;p&gt;Test and production traffic must be distinguishable. Record platform, application version, SDK version, and ingestion source where possible so failures can be isolated without inspecting individual users.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Run aggregate reconciliation queries
&lt;/h2&gt;

&lt;p&gt;After a single event is correct, send a small deterministic batch. For example, send ten events for three synthetic users with known properties. Then calculate the result directly in ClickHouse:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;event_count&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;uniqExact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;distinct_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;exact_users&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;sensors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'integration_test'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;properties&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'validation_run'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'2026-09-17-b'&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a validation batch, exact functions make expectations easier to reason about. Production queries may use other functions based on the required accuracy and cost; make that choice explicit in the metric definition.&lt;/p&gt;

&lt;p&gt;Also run negative checks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;countIf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;distinct_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;''&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;missing_identity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;countIf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event_time&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;missing_event_time&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;countIf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;received_at&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;event_time&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;INTERVAL&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;DAY&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;suspicious_future_time&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;sensors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;properties&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'validation_run'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'2026-09-17-b'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Adjust the conditions to the actual schema and business rules. The purpose is not to copy universal SQL, but to turn assumptions into executable tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: Reproduce the metric in Superset
&lt;/h2&gt;

&lt;p&gt;Connect Superset to the validated database using the project's documented driver and security configuration. Superset's &lt;a href="https://superset.apache.org/docs/databases/" rel="noopener noreferrer"&gt;database connection documentation&lt;/a&gt; lists supported connection patterns and points to database-specific requirements.&lt;/p&gt;

&lt;p&gt;Start in SQL Lab, run the same aggregate query, and compare the result with the ClickHouse client. Only after the results agree should you save a dataset and build a chart. Superset's &lt;a href="https://superset.apache.org/docs/using-superset/exploring-data" rel="noopener noreferrer"&gt;data exploration documentation&lt;/a&gt; explains the workflow from datasets to charts.&lt;/p&gt;

&lt;p&gt;When results differ, inspect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;dataset SQL or virtual-dataset transformations;&lt;/li&gt;
&lt;li&gt;dashboard and chart filters;&lt;/li&gt;
&lt;li&gt;time column selection and time grain;&lt;/li&gt;
&lt;li&gt;database and application time zones;&lt;/li&gt;
&lt;li&gt;cache state;&lt;/li&gt;
&lt;li&gt;row-level security and user permissions;&lt;/li&gt;
&lt;li&gt;distinct-count function and identity field;&lt;/li&gt;
&lt;li&gt;hidden test-environment filters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Save the validation query next to the metric definition. A screenshot alone is weak evidence because it omits filters, SQL, and data freshness.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 7: Validate retries, duplicates, and failure paths
&lt;/h2&gt;

&lt;p&gt;Happy-path testing is incomplete. Repeat the same request, interrupt the receiver during a small batch, and test malformed payloads in an isolated environment.&lt;/p&gt;

&lt;p&gt;Document the expected behavior for each case:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Expected behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Identical retry&lt;/td&gt;
&lt;td&gt;Stored once or explicitly identified as duplicate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Temporary database failure&lt;/td&gt;
&lt;td&gt;Retried with bounded backoff or placed in a recoverable queue&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Invalid required field&lt;/td&gt;
&lt;td&gt;Rejected with a diagnosable error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unsupported optional property&lt;/td&gt;
&lt;td&gt;Rejected or recorded according to documented policy, not silently lost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Receiver restart&lt;/td&gt;
&lt;td&gt;No unexplained accepted-but-missing events&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delayed mobile upload&lt;/td&gt;
&lt;td&gt;Original occurrence time retained and delay measurable&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The implementation may choose different guarantees, but operators need to know whether delivery is at-most-once, at-least-once, or effectively-once after deduplication.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 8: Use dual-write or sampled reconciliation during migration
&lt;/h2&gt;

&lt;p&gt;When replacing an existing receiver, avoid switching all traffic based on one successful test. Use dual-write, proxy mirroring, or a controlled traffic percentage where the client and infrastructure allow it.&lt;/p&gt;

&lt;p&gt;Compare both paths over fixed windows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;total events by event name;&lt;/li&gt;
&lt;li&gt;distinct identities using the same definition;&lt;/li&gt;
&lt;li&gt;property presence and type distribution;&lt;/li&gt;
&lt;li&gt;event-time delay distribution;&lt;/li&gt;
&lt;li&gt;duplicate rate;&lt;/li&gt;
&lt;li&gt;rejection and retry counts;&lt;/li&gt;
&lt;li&gt;a few business-critical funnel steps.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not expect every total to match without analysis. Different ingestion cutoffs, bot rules, late-event handling, identity resolution, and deduplication can produce legitimate differences. The goal is to explain the differences, not force cosmetic equality.&lt;/p&gt;

&lt;h2&gt;
  
  
  A release gate for event pipelines
&lt;/h2&gt;

&lt;p&gt;Before increasing traffic, require evidence for each item:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;[ ] A unique test event appears exactly as expected in raw storage.&lt;/li&gt;
&lt;li&gt;[ ] Identity transitions match the documented model.&lt;/li&gt;
&lt;li&gt;[ ] Client time and receipt time are both understood.&lt;/li&gt;
&lt;li&gt;[ ] Representative property types survive decoding and storage.&lt;/li&gt;
&lt;li&gt;[ ] A deterministic batch produces the expected event and user counts.&lt;/li&gt;
&lt;li&gt;[ ] ClickHouse and Superset return the same result for the validation query.&lt;/li&gt;
&lt;li&gt;[ ] Filters, time zone, cache, and permissions are recorded.&lt;/li&gt;
&lt;li&gt;[ ] Retry, duplicate, invalid payload, and restart behavior are tested.&lt;/li&gt;
&lt;li&gt;[ ] Rollback steps and owners are documented.&lt;/li&gt;
&lt;li&gt;[ ] Backups and restore procedures have been exercised for production data.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This checklist should be versioned with the deployment. Re-run it when the SDK, receiver, schema, database, driver, or dashboard definition changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where SensorFlow fits
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/data-analyze-bi/sensorFlow" rel="noopener noreferrer"&gt;SensorFlow&lt;/a&gt; provides a public Apache-2.0 implementation of the path used in this guide: compatible SDK event ingestion, a Go receiving service, ClickHouse storage, and Apache Superset for SQL exploration and dashboards. The repository exposes deployment files and code so teams can inspect the data path rather than relying only on marketing claims.&lt;/p&gt;

&lt;p&gt;It is best suited to engineering teams that want self-hosted data and are comfortable operating Docker, ClickHouse, and SQL-based analytics. It is not a drop-in replacement for every feature in a broad product analytics suite. Teams needing turnkey session replay, experimentation, feature flags, or extensive no-code analysis should compare more complete platforms as well.&lt;/p&gt;

&lt;p&gt;The detailed &lt;a href="https://sensorflow.site/use-cases/clickhouse-superset-analytics" rel="noopener noreferrer"&gt;ClickHouse and Superset implementation guide&lt;/a&gt; covers the architecture and operational boundaries. Teams migrating an existing compatible SDK path can also use the &lt;a href="https://sensorflow.site/use-cases/sensors-sdk-to-clickhouse" rel="noopener noreferrer"&gt;receiver-boundary migration guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Should an SDK write directly to ClickHouse?
&lt;/h3&gt;

&lt;p&gt;Generally, no. A receiving service should handle authentication, protocol decoding, validation, rate limits, retries, and error reporting. Exposing the database directly to untrusted clients expands security risk and makes protocol evolution harder.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is one matching event enough to approve a migration?
&lt;/h3&gt;

&lt;p&gt;No. One event validates the basic path. Approval should also cover a deterministic batch, identity transitions, property types, delayed events, duplicates, failures, aggregate SQL, and dashboard reconciliation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can Superset calculate funnels and retention?
&lt;/h3&gt;

&lt;p&gt;Superset can visualize SQL-derived funnel and retention results, but the team normally defines the model and query semantics. It does not automatically supply every specialized product-analytics workflow.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why keep both event time and receipt time?
&lt;/h3&gt;

&lt;p&gt;The two timestamps answer different questions. Event time represents when the action occurred on the client; receipt time shows when the platform observed it. Their difference helps identify offline uploads, network delays, clock problems, and processing backlogs.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the most important migration metric?
&lt;/h3&gt;

&lt;p&gt;There is no universal single metric. Start with unexplained event loss, identity consistency, critical-property validity, duplicate behavior, and agreement on business-critical counts. Choose thresholds before the traffic increase.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with one event, then prove the whole path
&lt;/h2&gt;

&lt;p&gt;Reliable analytics comes from traceability, not from a green endpoint alone. Create one identifiable event, inspect its raw row, verify its semantics, reconcile a known batch, and reproduce the query in Superset. Then test failure behavior and compare old and new paths under controlled traffic.&lt;/p&gt;

&lt;p&gt;To apply the checklist to a transparent reference implementation, review the &lt;a href="https://sensorflow.site/docs" rel="noopener noreferrer"&gt;SensorFlow deployment and validation documentation&lt;/a&gt; and inspect the source before sending production data.&lt;/p&gt;

</description>
      <category>clickhouse</category>
      <category>analytics</category>
      <category>opensource</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Choosing an Event Analytics Stack Without a Feature Checklist</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Mon, 14 Sep 2026 13:11:32 +0000</pubDate>
      <link>https://dev.to/sensorflow/choosing-an-event-analytics-stack-without-a-feature-checklist-3ch5</link>
      <guid>https://dev.to/sensorflow/choosing-an-event-analytics-stack-without-a-feature-checklist-3ch5</guid>
      <description>&lt;p&gt;Analytics comparisons often collapse into feature checklists. That is rarely how engineering teams experience the decision. The harder questions are migration cost, data ownership, metric governance, operational responsibility, and whether non-technical users need a purpose-built interface.&lt;/p&gt;

&lt;p&gt;This article compares categories rather than declaring a universal winner.&lt;/p&gt;

&lt;h2&gt;
  
  
  Full product analytics suites
&lt;/h2&gt;

&lt;p&gt;Sensors Analytics and Alibaba Cloud Quick Tracking represent the integrated commercial category. Public documentation shows broad collection, governance, and analysis workflows. Sensors Analytics, for example, documents coded tracking, visual auto-tracking, event analysis, funnels, retention, and paths.&lt;/p&gt;

&lt;p&gt;PostHog follows a similarly broad product approach internationally, connecting trends, funnels, retention, and paths with session replay, feature flags, and experiments. These suites fit teams that value ready-made workflows and want product or growth users to answer common questions without building every metric in SQL.&lt;/p&gt;

&lt;p&gt;The trade-offs to evaluate are procurement or hosting cost, migration, data residency, customization boundaries, and the operational model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lightweight web analytics
&lt;/h2&gt;

&lt;p&gt;Umami and Plausible are optimized for a different job: understandable, privacy-conscious website analytics. For teams mainly asking where traffic came from and which pages or goals perform well, a compact web analytics tool may be the best answer.&lt;/p&gt;

&lt;p&gt;That simplicity should not be confused with multi-platform product event analytics. Identity across web, mobile, mini-program, and backend events introduces a different set of requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open-source platforms with built-in analysis
&lt;/h2&gt;

&lt;p&gt;ClkLog presents a self-hosted platform with multi-platform collection, ClickHouse or Doris storage, and built-in analysis screens. Its public materials separate community capabilities from professional and CDP offerings, so teams should compare the exact edition while considering the Java, Kafka, and supporting infrastructure footprint.&lt;/p&gt;

&lt;h2&gt;
  
  
  SensorFlow's narrower boundary
&lt;/h2&gt;

&lt;p&gt;SensorFlow deliberately focuses on an inspectable data path:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Official Sensors Data SDK upload flow → Go ingestion → ClickHouse → Apache Superset&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For teams already using the standard upload flow of official Sensors Data SDKs, this creates a backend-first evaluation path. They can test ingestion, identity, property types, timestamps, and queries before deciding whether to replace every client integration.&lt;/p&gt;

&lt;p&gt;Raw events remain in ClickHouse controlled by the operator. Metrics can be expressed and reviewed in SQL, while Superset provides exploration, charts, and dashboards. This is useful when analytics must join internal business data or when teams want metric logic to remain inspectable.&lt;/p&gt;

&lt;p&gt;The boundary matters: SensorFlow is not a feature-for-feature replacement for PostHog, Sensors Analytics, or ClkLog. It does not currently claim mature no-code funnels, session replay, experimentation, feature flags, or visual auto-tracking. Self-hosting also makes the operator responsible for TLS, authentication, monitoring, backups, capacity, upgrades, and compliance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five questions to ask
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Is the goal website traffic reporting or multi-platform product events?&lt;/li&gt;
&lt;li&gt;Do product users need no-code analysis, or can the team work in SQL?&lt;/li&gt;
&lt;li&gt;Must raw events remain inside controlled infrastructure?&lt;/li&gt;
&lt;li&gt;How expensive would replacing existing client SDKs be?&lt;/li&gt;
&lt;li&gt;How much operational responsibility is acceptable in exchange for control?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Choose a lightweight tool when the problem is lightweight. Choose an integrated suite when workflow breadth matters most. Evaluate SensorFlow when the priority is a backend-first migration, ClickHouse ownership, and SQL-defined analysis.&lt;/p&gt;

&lt;p&gt;SensorFlow is an independent open-source project and is not affiliated with, endorsed by, or certified by the vendors mentioned here. Compatibility must be validated against actual SDK versions and event samples.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Source: &lt;a href="https://github.com/data-analyze-bi/sensorFlow" rel="noopener noreferrer"&gt;https://github.com/data-analyze-bi/sensorFlow&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Documentation: &lt;a href="https://sensorflow.site/docs" rel="noopener noreferrer"&gt;https://sensorflow.site/docs&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Read-only Superset demo: &lt;a href="https://superset.sensorflow.site/" rel="noopener noreferrer"&gt;https://superset.sensorflow.site/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>analytics</category>
      <category>selfhosted</category>
      <category>opensource</category>
      <category>clickhouse</category>
    </item>
    <item>
      <title>Moving product analytics safely: start with the ingestion boundary</title>
      <dc:creator>SensorFlow</dc:creator>
      <pubDate>Mon, 14 Sep 2026 07:09:09 +0000</pubDate>
      <link>https://dev.to/sensorflow/moving-product-analytics-safely-start-with-the-ingestion-boundary-33mp</link>
      <guid>https://dev.to/sensorflow/moving-product-analytics-safely-start-with-the-ingestion-boundary-33mp</guid>
      <description>&lt;p&gt;Replacing an analytics system is rarely a dashboard decision. The risky part is the event-ingestion boundary: web, Android, iOS, mini-program, and backend clients often emit events on different release cycles, with years of naming conventions and identity edge cases.&lt;/p&gt;

&lt;p&gt;SensorFlow is an open-source, self-hosted stack designed for a narrow, inspectable transition:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SDK event upload → Go ingestion service → ClickHouse → Apache Superset
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For teams already using the standard upload flow of official Sensors Data SDKs, the practical question becomes whether ingestion and storage can move into infrastructure they control without rewriting every client integration at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the stack makes visible
&lt;/h2&gt;

&lt;p&gt;SensorFlow processes incoming events in Go, stores them in ClickHouse, and makes them available to Apache Superset. That gives operators a direct path to inspect raw records, define metrics in SQL, and build dashboards against their own database.&lt;/p&gt;

&lt;p&gt;The point is not to hide analytics behind another hosted layer. It is to make event data operate like the rest of the data stack: reviewable schema changes, inspectable queries, and access controls that fit existing infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The operational boundary still matters
&lt;/h2&gt;

&lt;p&gt;Self-hosting is not a shortcut around operations. Docker Compose is appropriate for evaluation, while production requires HTTPS, unique credentials, access controls, monitoring, backups, restore tests, and capacity planning.&lt;/p&gt;

&lt;p&gt;SensorFlow also has deliberate scope limits. It is not a feature-for-feature replacement for broad product-analytics suites: session replay, experiments, feature flags, and broad no-code workflows are not provided today. It is intended for engineering teams with existing multi-platform instrumentation, a need for data ownership, and the willingness to operate ClickHouse.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evaluate the ingestion boundary
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/data-analyze-bi/sensorFlow" rel="noopener noreferrer"&gt;Open-source repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sensorflow.site/docs" rel="noopener noreferrer"&gt;Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://superset.sensorflow.site/" rel="noopener noreferrer"&gt;Read-only Superset demo&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I’m one of the maintainers. Feedback on ingestion compatibility, ClickHouse schema choices, deployment reliability, and production validation would be especially useful.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>dataengineering</category>
    </item>
  </channel>
</rss>
