<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Evgenii Timofeev</title>
    <description>The latest articles on DEV Community by Evgenii Timofeev (@eu_ti_f127c5b5d7535b7174f).</description>
    <link>https://dev.to/eu_ti_f127c5b5d7535b7174f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4102419%2Fd5567118-12e1-4443-add0-48855e37440c.jpg</url>
      <title>DEV Community: Evgenii Timofeev</title>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/eu_ti_f127c5b5d7535b7174f"/>
    <language>en</language>
    <item>
      <title>Google Sheets to Your Warehouse: the Share Step Is the One That Fails</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Sat, 26 Sep 2026 00:00:41 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/google-sheets-to-your-warehouse-the-share-step-is-the-one-that-fails-39m9</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/google-sheets-to-your-warehouse-the-share-step-is-the-one-that-fails-39m9</guid>
      <description>&lt;p&gt;Google Sheets is the most common shadow database in any company. Marketing keeps campaign spend in one, finance maintains the budget model in another, ops runs a lightweight CRM in a third. None of it is in the warehouse, so none of it can be joined to anything, so somebody copy-pastes it into a dashboard once a month and everyone quietly agrees not to ask how fresh it is.&lt;/p&gt;

&lt;p&gt;Landing a sheet in your warehouse is genuinely quick — a service account, four fields, a cron. This post is about the one step in the middle that fails for almost everybody the first time, and about a button that will refuse to tell you whether it worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup, briefly
&lt;/h2&gt;

&lt;p&gt;Google Sheets is &lt;strong&gt;source-only&lt;/strong&gt; in Datanika: you land it somewhere, you do not write back to it. You need a destination warehouse already connected — &lt;a href="https://datanika.io/connectors/postgresql/" rel="noopener noreferrer"&gt;PostgreSQL&lt;/a&gt; is the quickest if you do not have one yet — and a Google Cloud project with the &lt;strong&gt;Google Sheets API&lt;/strong&gt; enabled.&lt;/p&gt;

&lt;p&gt;If you already run &lt;a href="https://datanika.io/connectors/bigquery/" rel="noopener noreferrer"&gt;BigQuery&lt;/a&gt; with Datanika, reuse that same project and its service account; you only need to enable the Sheets API on it and do the sharing step below.&lt;/p&gt;

&lt;p&gt;Create a service account — &lt;strong&gt;IAM &amp;amp; Admin → Service accounts&lt;/strong&gt; — and give it &lt;strong&gt;no IAM roles at all&lt;/strong&gt;. This trips people up, because granting a role feels like the responsible thing to do. It is not relevant here: the service account reaches Sheets through the Sheets API, not through GCP resources, so a role grant would widen its access without enabling anything you need. Create a &lt;strong&gt;JSON key&lt;/strong&gt;, download it, and note the account's email address:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;datanika-sheets-reader@your-project.iam.gserviceaccount.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That address is the whole story of this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  The step that fails
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A Google service account can only read spreadsheets that have been explicitly shared with it.&lt;/strong&gt; Not sheets your account can see. Not sheets in a Drive folder it can reach. The specific spreadsheet, shared with that specific address, the same way you would share it with a colleague.&lt;/p&gt;

&lt;p&gt;So: open the sheet, click &lt;strong&gt;Share&lt;/strong&gt;, paste &lt;code&gt;datanika-sheets-reader@your-project.iam.gserviceaccount.com&lt;/code&gt;, set it to &lt;strong&gt;Viewer&lt;/strong&gt; — Datanika never writes to Google Sheets — and, the part that catches people, &lt;strong&gt;untick "Notify people"&lt;/strong&gt; before you confirm. The service account has no mailbox, so there is nobody to notify.&lt;/p&gt;

&lt;p&gt;Repeat for every spreadsheet you want this connection to read. There is no folder-level shortcut.&lt;/p&gt;

&lt;h2&gt;
  
  
  The button that says "not tested"
&lt;/h2&gt;

&lt;p&gt;Now add the connection. Open &lt;strong&gt;&lt;code&gt;/connections&lt;/code&gt;&lt;/strong&gt; — the form is already on the page, there is no "New Connection" button to hunt for — pick &lt;code&gt;google_sheets&lt;/code&gt; from the type dropdown, and fill in three fields:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Connection Name&lt;/strong&gt; — something you will recognise later, like &lt;code&gt;gsheets-marketing-budget&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spreadsheet URL&lt;/strong&gt; — the &lt;strong&gt;full&lt;/strong&gt; URL, &lt;code&gt;https://docs.google.com/spreadsheets/d/&amp;lt;ID&amp;gt;/edit&lt;/code&gt;, not just the ID.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Service Account JSON&lt;/strong&gt; — the entire contents of the key file. It is encrypted at rest with Fernet.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then click &lt;strong&gt;Test Connection&lt;/strong&gt;, and watch it decline to give you a verdict. It returns a neutral &lt;strong&gt;not tested&lt;/strong&gt;, carrying the reason.&lt;/p&gt;

&lt;p&gt;This is deliberate, and it is worth explaining, because a lot of tools would have shown you a green tick here.&lt;/p&gt;

&lt;p&gt;Verifying a service-account credential means minting an OAuth token. That is a real check, and it would pass — the JSON is well-formed, the key is valid, the API is enabled. But it would tell you nothing about the failure you are actually going to hit, because &lt;strong&gt;the sharing permission is checked per upload, not per connection&lt;/strong&gt;. A green tick at this point would be an accurate statement about the credential and a misleading one about the pipeline, and the person reading it cannot tell those apart.&lt;/p&gt;

&lt;p&gt;The alternative — a red cross — is worse in the other direction: a connection that is fine, reported as broken. Reporting an unverified thing as working and reporting it as failed are the same lie told in two directions. So the button says what is true: it has not been tested here, and &lt;strong&gt;the first real verification is the first run&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Click &lt;strong&gt;Create Connection&lt;/strong&gt; and go get one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Configure the upload
&lt;/h2&gt;

&lt;p&gt;Extract-load lives at &lt;strong&gt;&lt;code&gt;/uploads&lt;/code&gt;&lt;/strong&gt;, not on the connection. Connection rows offer Test / Edit / Copy / Delete and nothing else, and &lt;code&gt;/pipelines&lt;/code&gt; is the &lt;strong&gt;dbt&lt;/strong&gt; builder — a different thing entirely.&lt;/p&gt;

&lt;p&gt;Open &lt;code&gt;/uploads&lt;/code&gt;, and note the first quirk while you type: &lt;strong&gt;the upload name accepts letters, digits and spaces&lt;/strong&gt;, and strips everything else as you type. &lt;code&gt;sheets-daily-sync&lt;/code&gt; becomes &lt;code&gt;sheetsdailysync&lt;/code&gt; in front of you, while &lt;code&gt;Sheets Daily Sync&lt;/code&gt; is kept verbatim. This matters more than it looks, because the schedule you create later references the upload &lt;strong&gt;by name&lt;/strong&gt;, exactly as saved.&lt;/p&gt;

&lt;p&gt;Pick your source and destination connections — the pickers list entries as &lt;code&gt;16 — myconnection (postgres)&lt;/code&gt;, so id, name, type — and then set:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sheet Names&lt;/strong&gt; &lt;em&gt;(optional, comma-separated)&lt;/em&gt; — name the tabs you want. &lt;strong&gt;Leave it empty to load every tab in the spreadsheet.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batch size&lt;/strong&gt; — 10000 by default.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schema Contract&lt;/strong&gt; — three dropdowns, &lt;strong&gt;Tables&lt;/strong&gt; / &lt;strong&gt;Columns&lt;/strong&gt; / &lt;strong&gt;Data Type&lt;/strong&gt;, deciding whether a changed incoming shape evolves the destination or fails the run. For a spreadsheet that humans edit, this is the setting worth thinking about hardest: somebody &lt;em&gt;will&lt;/em&gt; add a column.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is no write disposition, load mode, source schema or table-name field, and that is not an omission — those controls are rendered only when the source is a SQL database.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first run is the actual test
&lt;/h2&gt;

&lt;p&gt;Hit &lt;strong&gt;Run&lt;/strong&gt; on the upload's row. There is no "Run now" on a pipeline page; the trigger lives on the upload's own row. Watch &lt;code&gt;/runs&lt;/code&gt;: status badge, start and finish timestamps, a &lt;strong&gt;Rows&lt;/strong&gt; count, and a &lt;strong&gt;Logs&lt;/strong&gt; icon for the detail.&lt;/p&gt;

&lt;p&gt;If the share step went wrong, this is where you find out.&lt;/p&gt;

&lt;p&gt;When it finishes, open &lt;strong&gt;Models&lt;/strong&gt; (&lt;code&gt;/models&lt;/code&gt;). Your tables land in a schema &lt;strong&gt;derived from the upload's name&lt;/strong&gt;: whitespace runs become single underscores and the whole thing is lower-cased. So &lt;code&gt;sheetsdailysync&lt;/code&gt; creates schema &lt;code&gt;sheetsdailysync&lt;/code&gt;, while &lt;code&gt;Sheets Daily Sync&lt;/code&gt; — the name this post told you is kept verbatim — creates schema &lt;code&gt;sheets_daily_sync&lt;/code&gt;. dlt also writes its own &lt;code&gt;_dlt_loads&lt;/code&gt;, &lt;code&gt;_dlt_pipeline_state&lt;/code&gt; and &lt;code&gt;_dlt_version&lt;/code&gt; bookkeeping tables into that schema, and Models does not list them; seeing only your own tables there is correct, not a partial load. There is no target-schema field to choose.&lt;/p&gt;

&lt;p&gt;Then do the thing the badge cannot do for you: &lt;strong&gt;spot-check the row count against the sheet.&lt;/strong&gt; A green run means the load finished. It does not mean it moved what you expected — an empty tab and a tab you forgot to share are very different problems that can produce the same colour.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put it on a cron
&lt;/h2&gt;

&lt;p&gt;Schedules are at &lt;strong&gt;&lt;code&gt;/schedules&lt;/code&gt;&lt;/strong&gt;, and reference the upload by name:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Target type&lt;/strong&gt; — &lt;code&gt;upload&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Target name&lt;/strong&gt; — exactly as saved, so &lt;code&gt;sheetsdailysync&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cron expression&lt;/strong&gt; — a real five-field cron string. There is no cadence picker and no "manual only" option; leaving an upload unscheduled &lt;em&gt;is&lt;/em&gt; manual-only. &lt;code&gt;0 * * * *&lt;/code&gt; hourly, &lt;code&gt;0 */6 * * *&lt;/code&gt; every six hours, &lt;code&gt;0 3 * * *&lt;/code&gt; nightly at 03:00.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Timezone&lt;/strong&gt; — &lt;code&gt;UTC&lt;/code&gt; by default, and the cron is evaluated in it, which matters for daily and weekly cadences.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then wire failure alerts in &lt;strong&gt;Settings → Notifications&lt;/strong&gt;, because a spreadsheet pipeline breaks for a reason no other pipeline has: somebody renames a tab.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-line version
&lt;/h2&gt;

&lt;p&gt;Everything in a Sheets sync is easy except remembering that the service account is a stranger to your spreadsheet until you introduce them. If a run fails and the credential is fine, you already know which step to go back to.&lt;/p&gt;

&lt;p&gt;Full field-by-field reference: the &lt;a href="https://datanika.io/docs/connectors/google-sheets/" rel="noopener noreferrer"&gt;Google Sheets setup guide&lt;/a&gt; and the &lt;a href="https://datanika.io/connectors/google-sheets/" rel="noopener noreferrer"&gt;connector page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>tutorial</category>
      <category>googlesheets</category>
      <category>connectors</category>
      <category>elt</category>
    </item>
    <item>
      <title>A Customer 360 from HubSpot and Stripe: the Join Is the Work</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Fri, 25 Sep 2026 08:07:46 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/a-customer-360-from-hubspot-and-stripe-the-join-is-the-work-22f1</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/a-customer-360-from-hubspot-and-stripe-the-join-is-the-work-22f1</guid>
      <description>&lt;p&gt;Every "customer 360" tutorial shows you the easy half. Sync your CRM, sync your billing system, write a &lt;code&gt;sources.yml&lt;/code&gt;, and then — with no ceremony at all — this line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="n"&gt;hubspot_contacts&lt;/span&gt; &lt;span class="n"&gt;hc&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;hc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That line is where the tutorial ends and the work starts. HubSpot and Stripe share &lt;strong&gt;no identifier&lt;/strong&gt;. Not one. There is no HubSpot ID in Stripe and no Stripe ID in HubSpot unless somebody deliberately put it there. The only overlapping &lt;em&gt;value&lt;/em&gt; is an email address, typed by a human into two different forms, possibly years apart, possibly not the same address.&lt;/p&gt;

&lt;p&gt;So the question this post answers is not "how do I join these." It is &lt;strong&gt;what fraction of your customers actually match, how do you find out, and what do you do with the ones that don't&lt;/strong&gt; — because in every real dataset there are some, and the difference between a trustworthy customer 360 and a misleading one is entirely in how honestly you handle them.&lt;/p&gt;

&lt;p&gt;If you want the Stripe-only revenue modelling first — MRR, churn, LTV — that is &lt;a href="https://datanika.io/blog/stripe-revenue-dashboard-dbt/" rel="noopener noreferrer"&gt;a separate post&lt;/a&gt;. This one assumes you have Stripe landed and adds HubSpot beside it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the two systems actually give you
&lt;/h2&gt;

&lt;p&gt;Worth being precise, because half the tutorials out there join against tables that don't exist.&lt;/p&gt;

&lt;p&gt;In Datanika, HubSpot and Stripe are SaaS sources, so the upload form shows &lt;strong&gt;Select endpoints to load&lt;/strong&gt; — a checkbox per resource. The full lists:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;Endpoints&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;HubSpot&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;companies&lt;/code&gt;, &lt;code&gt;contacts&lt;/code&gt;, &lt;code&gt;deals&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Stripe&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;charges&lt;/code&gt;, &lt;code&gt;customers&lt;/code&gt;, &lt;code&gt;invoices&lt;/code&gt;, &lt;code&gt;prices&lt;/code&gt;, &lt;code&gt;products&lt;/code&gt;, &lt;code&gt;subscriptions&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nine tables. Scan them for a shared key and you will not find one. &lt;code&gt;hubspot.contacts&lt;/code&gt; has an email and an associated company; &lt;code&gt;stripe.customers&lt;/code&gt; has an email and a name. That is the entire overlap, and both sides are free text.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where the tables land.&lt;/strong&gt; A SaaS source has no write-disposition, load-mode or target-schema field — those are rendered only for SQL database sources. On &lt;strong&gt;BigQuery / Snowflake / Databricks&lt;/strong&gt; the schema is the &lt;strong&gt;Dataset&lt;/strong&gt;/&lt;strong&gt;Schema&lt;/strong&gt; field on the &lt;em&gt;destination connection&lt;/em&gt;. On &lt;strong&gt;Postgres / MySQL / DuckDB&lt;/strong&gt; it is the &lt;strong&gt;upload's own name, snake-cased&lt;/strong&gt; — and upload names accept letters, digits and spaces only, so you name the upload &lt;strong&gt;&lt;code&gt;Raw Hubspot&lt;/code&gt;&lt;/strong&gt; and get the schema &lt;code&gt;raw_hubspot&lt;/code&gt;. Check the &lt;strong&gt;Schema&lt;/strong&gt; column in &lt;strong&gt;Models&lt;/strong&gt; — the sidebar entry for Datanika's data catalog, at &lt;strong&gt;&lt;code&gt;/models&lt;/code&gt;&lt;/strong&gt; — for the real name before writing &lt;code&gt;sources.yml&lt;/code&gt;; a wrong schema name is the most common reason the models below don't resolve.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Two uploads, two schemas:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# models/staging/sources.yml&lt;/span&gt;
&lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;

&lt;span class="na"&gt;sources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;raw_stripe&lt;/span&gt;
    &lt;span class="na"&gt;tables&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[{&lt;/span&gt;&lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;customers&lt;/span&gt;&lt;span class="pi"&gt;},&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;subscriptions&lt;/span&gt;&lt;span class="pi"&gt;},&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;invoices&lt;/span&gt;&lt;span class="pi"&gt;}]&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;raw_hubspot&lt;/span&gt;
    &lt;span class="na"&gt;tables&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[{&lt;/span&gt;&lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;contacts&lt;/span&gt;&lt;span class="pi"&gt;},&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;companies&lt;/span&gt;&lt;span class="pi"&gt;},&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;deals&lt;/span&gt;&lt;span class="pi"&gt;}]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;HubSpot returns default properties only.&lt;/strong&gt; If &lt;code&gt;email&lt;/code&gt; or &lt;code&gt;domain&lt;/code&gt; is missing from a landed table, it is not a Datanika bug — HubSpot's API returns a default property set unless the pipeline asks for more. Add the custom properties you need to the resource config before you go hunting through dbt.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Step 1 — Normalize the identifiers before you join anything
&lt;/h2&gt;

&lt;p&gt;Never join on a raw email column. Email is case-insensitive in the domain part, effectively case-insensitive at every provider anyone uses, frequently has trailing whitespace from a paste, and Gmail ignores dots and &lt;code&gt;+tags&lt;/code&gt;. Two staging models, one job each: produce a &lt;code&gt;match_email&lt;/code&gt; you can trust and a &lt;code&gt;email_domain&lt;/code&gt; you can fall back to.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/staging/stg_hubspot__contacts.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'view'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;hs_object_id&lt;/span&gt;                              &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;contact_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;associatedcompanyid&lt;/span&gt;                       &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;company_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;email&lt;/span&gt;                                     &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;email_raw&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;                        &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;match_email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt; &lt;span class="s1"&gt;'@'&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="k"&gt;offset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;email_domain&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;-- BigQuery&lt;/span&gt;
    &lt;span class="n"&gt;firstname&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;lastname&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;createdate&lt;/span&gt;                                &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'raw_hubspot'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'contacts'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/staging/stg_stripe__customers.sql  (extends the version in the revenue post)&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'view'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt;                                        &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;email&lt;/span&gt;                                     &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;email_raw&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;                        &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;match_email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt; &lt;span class="s1"&gt;'@'&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="k"&gt;offset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;email_domain&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="c1"&gt;-- the exact join key, when someone was disciplined enough to write it&lt;/span&gt;
    &lt;span class="n"&gt;json_value&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'$.hubspot_contact_id'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;hubspot_contact_id_from_metadata&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timestamp_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;created&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'raw_stripe'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'customers'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Dialect.&lt;/strong&gt; &lt;code&gt;split(...)[offset(1)]&lt;/code&gt; and &lt;code&gt;json_value&lt;/code&gt; are BigQuery. Snowflake: &lt;code&gt;split_part(lower(trim(email)), '@', 2)&lt;/code&gt; and &lt;code&gt;metadata:hubspot_contact_id::string&lt;/code&gt;. Postgres/DuckDB: &lt;code&gt;split_part(...)&lt;/code&gt; and &lt;code&gt;metadata -&amp;gt;&amp;gt; 'hubspot_contact_id'&lt;/code&gt;. Everything after staging is plain SQL.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The two remaining HubSpot models are thin. Note that &lt;code&gt;companies.domain&lt;/code&gt; gets the same &lt;code&gt;lower(trim(...))&lt;/code&gt; treatment — it is the Tier C join key, and a stray capital there costs you matches for no reason:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/staging/stg_hubspot__companies.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'view'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;hs_object_id&lt;/span&gt;        &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;company_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;domain&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="k"&gt;domain&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'raw_hubspot'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'companies'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="k"&gt;domain&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/staging/stg_hubspot__deals.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'view'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;hs_object_id&lt;/span&gt;       &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;deal_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;associatedcompanyid&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;company_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;-- see the caveat below&lt;/span&gt;
    &lt;span class="n"&gt;dealname&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;dealstage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;closedate&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'raw_hubspot'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'deals'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Check the deal→company column before you trust it.&lt;/strong&gt; In HubSpot, a deal's link to a company is an &lt;em&gt;association&lt;/em&gt;, not a plain property, and which column (if any) lands depends on the properties your pipeline requests. Open &lt;strong&gt;Models&lt;/strong&gt; and look at the landed &lt;code&gt;deals&lt;/code&gt; table. If there is no company column, either add the association to the resource config or drop the two deal columns from the final model — the customer 360 is still worth building without them. Do not join deals on company &lt;em&gt;name&lt;/em&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Step 2 — Three joins, in descending order of trust
&lt;/h2&gt;

&lt;p&gt;This is the part worth reading twice. There is not one join. There are three, they have very different reliability, and a customer 360 that treats them as interchangeable is quietly wrong.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tier A — an explicit key in Stripe &lt;code&gt;metadata&lt;/code&gt; (exact)
&lt;/h3&gt;

&lt;p&gt;Stripe lets you attach arbitrary &lt;code&gt;metadata&lt;/code&gt; to a customer. If your signup flow writes the HubSpot contact ID there, you have a real foreign key and none of the rest of this post applies to those rows.&lt;/p&gt;

&lt;p&gt;Almost nobody has this on day one, because it has to be decided &lt;em&gt;before&lt;/em&gt; the customers exist. &lt;strong&gt;If you take one action from this post, make it this one&lt;/strong&gt;: start writing &lt;code&gt;hubspot_contact_id&lt;/code&gt; into Stripe customer metadata today. It does nothing for your backlog and it makes every future row exact.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tier B — normalized email (good, and incomplete)
&lt;/h3&gt;

&lt;p&gt;The workhorse. Matches whenever the human typed the same address into both systems.&lt;/p&gt;

&lt;p&gt;It fails in ways worth knowing, because each one is a real customer in your unmatched pile:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;signed up with &lt;code&gt;dana@acme.com&lt;/code&gt;, billing goes through &lt;code&gt;accounts-payable@acme.com&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;signed up with a personal address, expensed it later&lt;/li&gt;
&lt;li&gt;the company was acquired and the domain changed&lt;/li&gt;
&lt;li&gt;one side has a typo&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Tier C — email domain → company domain (recovers real customers, over-matches on public domains)
&lt;/h3&gt;

&lt;p&gt;When the emails differ but both live at &lt;code&gt;acme.com&lt;/code&gt;, HubSpot's &lt;code&gt;companies.domain&lt;/code&gt; bridges them. This is how you catch the &lt;code&gt;accounts-payable@&lt;/code&gt; case, and it is the tier that needs a guard:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;gmail.com&lt;/code&gt; is not a company.&lt;/strong&gt; Join two consumer addresses on their domain and you will merge unrelated people into one "customer" with a straight face. Exclude free providers explicitly, and treat a domain matching many distinct Stripe customers as a signal to stop rather than a bigger match.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/marts/customer_identity_map.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'table'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;stripe&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_stripe__customers'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}),&lt;/span&gt;
     &lt;span class="n"&gt;contacts&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_hubspot__contacts'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}),&lt;/span&gt;
     &lt;span class="n"&gt;companies&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_hubspot__companies'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}),&lt;/span&gt;

&lt;span class="n"&gt;free_domains&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="k"&gt;domain&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="k"&gt;unnest&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
        &lt;span class="s1"&gt;'gmail.com'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'googlemail.com'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'yahoo.com'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'outlook.com'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'hotmail.com'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s1"&gt;'live.com'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'icloud.com'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'me.com'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'proton.me'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'protonmail.com'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'aol.com'&lt;/span&gt;
    &lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="k"&gt;domain&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;

&lt;span class="c1"&gt;-- Tier C is only trustworthy where the domain identifies ONE company. A domain&lt;/span&gt;
&lt;span class="c1"&gt;-- shared by many Stripe customers is either a free provider we missed or an&lt;/span&gt;
&lt;span class="c1"&gt;-- agency billing for several clients; both should fall through to unmatched.&lt;/span&gt;
&lt;span class="n"&gt;safe_domains&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;email_domain&lt;/span&gt;
    &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;stripe&lt;/span&gt;
    &lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;email_domain&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="k"&gt;domain&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;free_domains&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;group&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="n"&gt;email_domain&lt;/span&gt;
    &lt;span class="k"&gt;having&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;distinct&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;

&lt;span class="n"&gt;tier_a&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contact_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'A_metadata'&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;match_tier&lt;/span&gt;
    &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;stripe&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
    &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="n"&gt;contacts&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contact_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;hubspot_contact_id_from_metadata&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;

&lt;span class="n"&gt;tier_b&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contact_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'B_email'&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;match_tier&lt;/span&gt;
    &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;stripe&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
    &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="n"&gt;contacts&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;match_email&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;match_email&lt;/span&gt;
    &lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tier_a&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;

&lt;span class="n"&gt;tier_c&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contact_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'C_domain'&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;match_tier&lt;/span&gt;
    &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;stripe&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
    &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="n"&gt;safe_domains&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email_domain&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email_domain&lt;/span&gt;
    &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="n"&gt;companies&lt;/span&gt; &lt;span class="n"&gt;co&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;co&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;domain&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email_domain&lt;/span&gt;
    &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="n"&gt;contacts&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;company_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;co&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;company_id&lt;/span&gt;
    &lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tier_a&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tier_b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;

&lt;span class="n"&gt;unmatched&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;cast&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;null&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;contact_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'UNMATCHED'&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;match_tier&lt;/span&gt;
    &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;stripe&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
    &lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tier_a&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tier_b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tier_c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tier_a&lt;/span&gt;
&lt;span class="k"&gt;union&lt;/span&gt; &lt;span class="k"&gt;all&lt;/span&gt; &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tier_b&lt;/span&gt;
&lt;span class="k"&gt;union&lt;/span&gt; &lt;span class="k"&gt;all&lt;/span&gt; &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tier_c&lt;/span&gt;
&lt;span class="k"&gt;union&lt;/span&gt; &lt;span class="k"&gt;all&lt;/span&gt; &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;unmatched&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two design choices worth stating, because they are the difference between this and the one-liner:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Every Stripe customer appears exactly once&lt;/strong&gt;, including the ones that matched nothing. A map that silently drops rows turns "we couldn't match 18% of your revenue" into "revenue is 18% lower than Stripe says," and someone will spend a day on that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;match_tier&lt;/code&gt; is a column, not a comment.&lt;/strong&gt; It travels downstream, so any number built on this map can be recomputed for exact matches only.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Step 3 — Measure the match rate; don't assume it
&lt;/h2&gt;

&lt;p&gt;Here is the model that makes this a system instead of a query. It answers &lt;em&gt;"how much of this do I believe?"&lt;/em&gt; and it answers it every run.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/marts/customer_identity_coverage.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'table'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;match_tier&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                                                    &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="n"&gt;over&lt;/span&gt; &lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;          &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pct_of_customers&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'customer_identity_map'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;span class="k"&gt;group&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="n"&gt;match_tier&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then turn the number into a gate. dbt tests fail builds; that is the point:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# models/marts/schema.yml&lt;/span&gt;
&lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;

&lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;customer_identity_map&lt;/span&gt;
    &lt;span class="na"&gt;columns&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;customer_id&lt;/span&gt;
        &lt;span class="na"&gt;tests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;not_null&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;unique&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;        &lt;span class="c1"&gt;# a fan-out here means a duplicate contact&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;match_tier&lt;/span&gt;
        &lt;span class="na"&gt;tests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;accepted_values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;A_metadata'&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;B_email'&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;C_domain'&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;UNMATCHED'&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- tests/assert_identity_coverage_above_threshold.sql&lt;/span&gt;
&lt;span class="c1"&gt;-- Fails when unmatched customers exceed 25%. Set the threshold to a little worse&lt;/span&gt;
&lt;span class="c1"&gt;-- than today's real number, then tighten it as you fix the data. A threshold you&lt;/span&gt;
&lt;span class="c1"&gt;-- have never been near is not a test — it is decoration.&lt;/span&gt;
&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;select&lt;/span&gt;
        &lt;span class="n"&gt;countif&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;match_tier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'UNMATCHED'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;unmatched_rate&lt;/span&gt;
    &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'customer_identity_map'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;unmatched_rate&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;25&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it the first time and read the number before you set the threshold. Whatever it is, it is &lt;em&gt;your&lt;/em&gt; number, and it is the most useful single fact in this entire pipeline: nobody who joins on email has any idea what theirs is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4 — Now the 360 is boring, which is the goal
&lt;/h2&gt;

&lt;p&gt;With a map that every row goes through, the customer view is a straightforward assembly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/marts/customer_360.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'table'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;match_tier&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;sc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email_raw&lt;/span&gt;                    &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;billing_email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;sc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;                         &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;billing_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;hc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;firstname&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lastname&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;co&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;                         &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;company_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;co&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;domain&lt;/span&gt;                       &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;company_domain&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="c1"&gt;-- billing side&lt;/span&gt;
    &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;distinct&lt;/span&gt; &lt;span class="n"&gt;sub&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;subscription_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                          &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;subscriptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;countif&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sub&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'active'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                               &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;active_subscriptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;coalesce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inv&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;amount_paid&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                            &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;lifetime_revenue&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                                           &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;billing_since&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="c1"&gt;-- CRM side&lt;/span&gt;
    &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;distinct&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deal_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                                    &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;deals&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;closedate&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                                             &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;last_deal_closed_at&lt;/span&gt;

&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'customer_identity_map'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;
&lt;span class="k"&gt;left&lt;/span&gt; &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_stripe__customers'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;    &lt;span class="n"&gt;sc&lt;/span&gt; &lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;left&lt;/span&gt; &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_hubspot__contacts'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;    &lt;span class="n"&gt;hc&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;hc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contact_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contact_id&lt;/span&gt;
&lt;span class="k"&gt;left&lt;/span&gt; &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_hubspot__companies'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;   &lt;span class="n"&gt;co&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;co&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;company_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;company_id&lt;/span&gt;
&lt;span class="k"&gt;left&lt;/span&gt; &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_stripe__subscriptions'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="n"&gt;sub&lt;/span&gt; &lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;left&lt;/span&gt; &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_stripe__invoices'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;     &lt;span class="n"&gt;inv&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;inv&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;
                                                   &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="n"&gt;inv&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'paid'&lt;/span&gt;
&lt;span class="k"&gt;left&lt;/span&gt; &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_hubspot__deals'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;       &lt;span class="n"&gt;d&lt;/span&gt;  &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;company_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;co&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;company_id&lt;/span&gt;
&lt;span class="k"&gt;group&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every join is a &lt;code&gt;left join&lt;/code&gt; &lt;strong&gt;from the map&lt;/strong&gt;, so an unmatched Stripe customer still produces a row with real billing figures and null CRM fields. That is the honest shape: you know what you know.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5 — Ship the unmatched list to the humans
&lt;/h2&gt;

&lt;p&gt;The unmatched bucket is not an error state. It is a work queue, and it is the highest-value output of the whole build — every row is a paying customer your CRM cannot see.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/marts/unmatched_paying_customers.sql&lt;/span&gt;
&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;billing_email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;billing_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lifetime_revenue&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'customer_360'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;match_tier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'UNMATCHED'&lt;/span&gt;
  &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lifetime_revenue&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="k"&gt;order&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lifetime_revenue&lt;/span&gt; &lt;span class="k"&gt;desc&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sorted by revenue, because the ten unmatched customers who pay you the most are worth an afternoon of manual reconciliation and the long tail is worth an automation. Every row someone fixes moves permanently into Tier A if they write the metadata back.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It is not identity resolution as a product.&lt;/strong&gt; No fuzzy name matching, no phonetic keys, no probabilistic scoring. Deterministic tiers with an explicit unmatched bucket beat a similarity threshold you cannot explain to finance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It does not deduplicate within a system.&lt;/strong&gt; Two HubSpot contacts for the same person will both match the same Stripe customer and trip the &lt;code&gt;unique&lt;/code&gt; test on &lt;code&gt;customer_id&lt;/code&gt; — deliberately. Fix the CRM; don't paper over it in SQL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It has no opinion on which side is right&lt;/strong&gt; when HubSpot and Stripe disagree about a name or a company. The model keeps both columns and lets you choose.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tier C is a heuristic.&lt;/strong&gt; The &lt;code&gt;count(distinct customer_id) = 1&lt;/code&gt; guard and the free-domain list make it defensible, not correct. If a number would embarrass you at 5% error, compute it on Tier A and B only — &lt;code&gt;match_tier&lt;/code&gt; is right there.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Wire it up
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Two connections — &lt;a href="https://datanika.io/docs/connectors/hubspot/" rel="noopener noreferrer"&gt;HubSpot&lt;/a&gt; (private app token, &lt;code&gt;crm.objects.{contacts,companies,deals}.read&lt;/code&gt;) and &lt;a href="https://datanika.io/docs/connectors/stripe/" rel="noopener noreferrer"&gt;Stripe&lt;/a&gt; (a &lt;strong&gt;restricted&lt;/strong&gt; read key, never a secret key).&lt;/li&gt;
&lt;li&gt;Two uploads. Name them so the schemas come out right, and confirm in &lt;strong&gt;Models&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The models above under &lt;strong&gt;Transformations&lt;/strong&gt;, then schedule the transform to depend on &lt;strong&gt;both&lt;/strong&gt; uploads — a 360 built from a fresh Stripe and yesterday's HubSpot is a 360 that disagrees with itself. See the &lt;a href="https://datanika.io/docs/scheduling/" rel="noopener noreferrer"&gt;scheduling guide&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Point a dashboard at &lt;code&gt;customer_360&lt;/code&gt; and &lt;code&gt;customer_identity_coverage&lt;/code&gt;. Put the coverage number on the dashboard, not in a runbook. A match rate nobody looks at drifts.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Datanika meters &lt;strong&gt;bytes processed&lt;/strong&gt;; both of these are narrow JSON sources and the &lt;a href="https://datanika.io/pricing/" rel="noopener noreferrer"&gt;Free plan&lt;/a&gt; includes 10 GB/month. Since 2026-09-24 the dashboard's &lt;strong&gt;Plan Usage&lt;/strong&gt; panel carries a &lt;strong&gt;bytes processed&lt;/strong&gt; dimension against your plan's included volume (&lt;a href="https://github.com/datanika-io/datanika-core/issues/1513" rel="noopener noreferrer"&gt;datanika-core#1513&lt;/a&gt;); it appears once a pipeline has written volume data, so there is nothing to read there before your first run. Size the first estimate from your own data — the meter counts what an upload writes after normalization.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next steps
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/blog/stripe-revenue-dashboard-dbt/" rel="noopener noreferrer"&gt;Stripe revenue dashboard with dbt&lt;/a&gt;&lt;/strong&gt; — MRR, churn and LTV on the Stripe half.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/docs/connectors/hubspot/" rel="noopener noreferrer"&gt;HubSpot setup guide&lt;/a&gt;&lt;/strong&gt; · &lt;strong&gt;&lt;a href="https://datanika.io/docs/connectors/stripe/" rel="noopener noreferrer"&gt;Stripe setup guide&lt;/a&gt;&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/docs/transformations/" rel="noopener noreferrer"&gt;Transformations&lt;/a&gt;&lt;/strong&gt; — how models, tests and schedules fit together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/connectors/" rel="noopener noreferrer"&gt;All connectors&lt;/a&gt;&lt;/strong&gt; — the same identity map takes a third source the day you add one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The join is the work. Measure it, publish the number, and hand the leftovers to a human.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Correction, 2026-09-23.&lt;/em&gt; The cost line told you to check &lt;strong&gt;Usage&lt;/strong&gt; for your own byte figures. That card counted model runs, not bytes, and at the time no screen in the app showed a byte count, so the line was changed to size it from your data instead. The same sentence was corrected in &lt;a href="https://datanika.io/blog/stripe-revenue-dashboard-dbt/" rel="noopener noreferrer"&gt;the Stripe revenue post&lt;/a&gt; a day earlier; tracked in &lt;a href="https://github.com/datanika-io/datanika-core/issues/1513" rel="noopener noreferrer"&gt;datanika-core#1513&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Update, 2026-09-24.&lt;/em&gt; &lt;a href="https://github.com/datanika-io/datanika-core/issues/1513" rel="noopener noreferrer"&gt;datanika-core#1513&lt;/a&gt; shipped, and the dashboard's usage card now carries a bytes-processed dimension — verified on the serving container across that day's deploy. The cost line names the screen again, together with the condition that makes the figure appear.&lt;/p&gt;

</description>
      <category>tutorial</category>
      <category>dbt</category>
      <category>hubspot</category>
      <category>stripe</category>
    </item>
    <item>
      <title>We're Adding a Volume Dimension to Our Pricing. Here's the Math, and Here's Why.</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Thu, 24 Sep 2026 09:01:47 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/were-adding-a-volume-dimension-to-our-pricing-heres-the-math-and-heres-why-42ik</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/were-adding-a-volume-dimension-to-our-pricing-heres-the-math-and-heres-why-42ik</guid>
      <description>&lt;p&gt;Our v1 pricing had a known hole. We're closing it now.&lt;/p&gt;

&lt;p&gt;Specifically: we charged $79/mo for "Pro" with a 15,000-runs-per-month quota and no cap on the &lt;em&gt;volume&lt;/em&gt; of data those runs moved. On paper, fine. In practice, a single customer running one pipeline that shovels 1 TB of Postgres into BigQuery per month would cost us $50–$150 in infrastructure to serve — on a $79 bill. Repeat that across ten customers and the business is upside-down.&lt;/p&gt;

&lt;p&gt;We don't have ten of those customers yet. We don't have &lt;em&gt;one&lt;/em&gt; of them yet. Which is exactly why we're fixing this now, before the first one arrives — not after, when it turns into an apology post.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;Old&lt;/th&gt;
&lt;th&gt;New&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Free&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;500 runs, no volume cap&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;10 GB processed/mo&lt;/strong&gt; + 500 runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Pro&lt;/strong&gt; ($79/mo)&lt;/td&gt;
&lt;td&gt;15,000 runs, no volume cap&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;100 GB processed/mo&lt;/strong&gt;, &lt;strong&gt;$0.50/extra GB&lt;/strong&gt;, 15,000 runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Enterprise&lt;/strong&gt; (from $399/mo)&lt;/td&gt;
&lt;td&gt;50,000 runs, $0.01/run overage, no volume cap&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1 TB processed/mo&lt;/strong&gt;, &lt;strong&gt;$0.25/extra GB&lt;/strong&gt;, 50,000 runs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Seats, connections, and schedules stay the same. SSO stays on Enterprise. The tiers haven't moved; the meter has.&lt;/p&gt;

&lt;p&gt;Every GB on this page — and every number derived from one below — is the &lt;strong&gt;binary&lt;/strong&gt; GB our meter counts: &lt;strong&gt;1,073,741,824 bytes (2&lt;sup&gt;30&lt;/sup&gt;)&lt;/strong&gt;, which is 7.4% more data than the decimal GB most warehouse consoles report. Worth knowing before you check our arithmetic against your own, because counting in decimal makes your usage look larger than we meter it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "volume" was missing
&lt;/h2&gt;

&lt;p&gt;Because we shipped v1 with Paddle's default-shape subscription plans — a flat monthly fee plus a secondary usage meter (model runs). That captured the obvious cost driver (orchestration, scheduler CPU, log storage) but ignored the expensive one (the actual bytes that hit disk, get normalized, get re-read by dbt, and get written to the destination).&lt;/p&gt;

&lt;p&gt;Run counts correlate with &lt;em&gt;some&lt;/em&gt; costs. They don't correlate with the cost that scales with your success as a customer. A pipeline running 30×/day with 500 rows per run is cheap for us. A pipeline running once a day with a full-history dump of your production database is not. v1 charged both the same.&lt;/p&gt;

&lt;p&gt;That's the hole.&lt;/p&gt;

&lt;h2&gt;
  
  
  The math that broke it
&lt;/h2&gt;

&lt;p&gt;Take one customer running one pipeline: a nightly export of a Postgres table with a million wide rows. Each row is ~1 KB, so the raw export is ~1 GB. Nested JSON columns flatten to ~3 GB after normalization. A dbt model aggregates that to a 100 MB summary. Total bytes touched per run: roughly 3.1 GB. Repeat nightly for 30 days: ~93 GB/mo.&lt;/p&gt;

&lt;p&gt;At v1, we billed $79 for this. At cloud-provider rates for storage + CPU + egress, we spent $12–$25 on infrastructure. That's fine. We're profitable on this shape.&lt;/p&gt;

&lt;p&gt;Now scale it to the same customer adding three more pipelines: a Stripe export, a HubSpot sync, a Segment event feed. Each one amplifies 2–5× after normalization. Total bytes/mo: ~400–600 GB. Infrastructure cost: $50–$100. Still $79 on the bill.&lt;/p&gt;

&lt;p&gt;Scale once more to a customer with a real data footprint — 1 TB/mo across 8 pipelines — and we're spending $120+ on a $79 subscription. That's the bill that convinced us the pricing was wrong.&lt;/p&gt;

&lt;p&gt;If you're processing more than &lt;strong&gt;740 GB/mo&lt;/strong&gt; — binary GB, as above; about 795 GB as a warehouse console counts them — Enterprise's $0.25/GB rate saves you more than the subscription difference. The &lt;a href="https://datanika.io/why-cheaper/" rel="noopener noreferrer"&gt;pricing calculator&lt;/a&gt; auto-picks the cheaper tier for you — no mental math required.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why GB and not MAR
&lt;/h2&gt;

&lt;p&gt;Fivetran's "monthly active rows" pricing is the obvious alternative. We looked at it; we're not doing it.&lt;/p&gt;

&lt;p&gt;MAR punishes schema choices you didn't make. A table with 10M narrow rows and a table with 100K wide rows can hit an identical disk-and-CPU bill but a 100× MAR bill. MAR also re-counts edits — update one row ten times this month and Fivetran charges you for ten rows. The customer has no way to predict the bill until the sync runs.&lt;/p&gt;

&lt;p&gt;GB measures the cost driver directly. A wide JSON blob and a narrow normalized row land on the same gram of disk for the same price. Updates don't inflate the count. You can predict your bill by looking at &lt;code&gt;du -sh&lt;/code&gt; on your source and multiplying by a constant.&lt;/p&gt;

&lt;p&gt;Fivetran publishes no per-MAR rate, so its side of this comparison is an illustrative estimate rather than a quote. Converting at ~200K rows/GB, 100 GB on Starter lands in the region of &lt;strong&gt;$3,800/mo&lt;/strong&gt; and 1 TB somewhere around &lt;strong&gt;$22,000/mo&lt;/strong&gt;. On Datanika Pro at 100 GB you pay $79, and on Enterprise at 1 TB, $399 + nothing (it's included) — those two are exact. The &lt;a href="https://datanika.io/why-cheaper/" rel="noopener noreferrer"&gt;calculator on /why-cheaper/&lt;/a&gt; shows the comparison side-by-side with a slider, and &lt;a href="https://www.fivetran.com/pricing" rel="noopener noreferrer"&gt;Fivetran's own estimator&lt;/a&gt; is where a binding number comes from.&lt;/p&gt;

&lt;p&gt;That's what we mean by "you pay for bytes, not tables."&lt;/p&gt;

&lt;h2&gt;
  
  
  How we meter honestly
&lt;/h2&gt;

&lt;p&gt;We count &lt;strong&gt;output bytes after normalization&lt;/strong&gt; — the amplified number, not the raw input. We do this because the amplified number is what our infrastructure actually touches, and pretending otherwise creates a gap between the sticker and the bill.&lt;/p&gt;

&lt;p&gt;Worked example: you have a 1 GB HubSpot JSON export. Our ingestion flattens nested objects into a wide table — that's ~3 GB of post-normalization data, and &lt;strong&gt;~3 GB is what counts against your quota&lt;/strong&gt;, not 1 GB. A dbt model that aggregates it to a 100 MB summary adds nothing to that: only uploads are metered in bytes, and a model run counts as a model run.&lt;/p&gt;

&lt;p&gt;The reason we publish the amplification rather than hiding it is that it makes the bill something you can work out in advance yourself. A gigabyte is a unit you can count; the rate is on the pricing page; the multiplication is yours. MAR is the opposite kind of unit — it is defined and counted by the vendor, so the invoice is the first place you get to see it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a month costs
&lt;/h2&gt;

&lt;p&gt;The meter reads &lt;em&gt;what an upload wrote&lt;/em&gt;: your data is normalized on our side, so a 1 GB JSON export becomes ~3 GB of flat tables and the meter counts ~3 GB.&lt;/p&gt;

&lt;p&gt;At $0.50/GB overage (Pro), each nightly run past your included 100 GB adds 3 GB to the month's metered volume — $1.50 worth. Over 30 nightly runs that's 90 GB: &lt;strong&gt;$45&lt;/strong&gt;. (Overage is totalled once per billing cycle and rounded up to the next whole GB, so that per-run figure is the monthly bill decomposed — not a charge you're billed run by run.)&lt;/p&gt;

&lt;h2&gt;
  
  
  What stays
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Self-host is still $0 forever.&lt;/strong&gt; The AGPL-3.0 open-source core has zero pricing dimensions. Run it on your own hardware, ingest 10 TB/mo, pay us nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;10 GB Free is real headroom.&lt;/strong&gt; That's enough for a real side project or a 3-source trickle-volume evaluation stack — a genuine test with production-shaped data, not a crippled sandbox. Fivetran Free tops out at 500K MAR (~1–2 GB equivalent); Hevo Free at 1M events. We offer 5–10× more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;All 36 connectors on every tier.&lt;/strong&gt; Free users don't get a crippled connector list. Fivetran adds a $5 base charge to each standard connection using under 1M MAR a month; we don't, and we're not going to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pro's 5 seats, Enterprise's 10 seats, schedules unlimited on paid tiers.&lt;/strong&gt; Seat economics are unchanged.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Annual discount at 17%&lt;/strong&gt; (Pro $79 → $66/mo billed annually; Enterprise from $399 → $333). We'll revisit after 90 days of real signup data — if the math says 20% works, we'll adjust and blog about it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What we'll do if we got the numbers wrong
&lt;/h2&gt;

&lt;p&gt;We picked 10 GB / 100 GB / 1 TB based on the cost model in &lt;a href="https://datanika.io/blog/real-cost-modern-data-stack/" rel="noopener noreferrer"&gt;price_insights.md §7&lt;/a&gt; and a napkin-margin target of 80–95% on variable cost. We don't yet have real signup data to validate those numbers against customer reality.&lt;/p&gt;

&lt;p&gt;Commitment: 90 days from now we look at actual usage distributions against actual infrastructure cost. If 100 GB on Pro is the wrong number — too tight for the real median customer, or too generous for our margin — we adjust the Pro tier's included volume. We blog about it when we do. The &lt;em&gt;overage rate&lt;/em&gt; won't move in year one; we picked $0.50/GB and $0.25/GB to sit comfortably above our variable cost at every volume we've modeled.&lt;/p&gt;

&lt;p&gt;If you're already on Datanika when we adjust: your current tier honors its current numbers until your next renewal, and we email you the change at least 30 days before it hits. Not that there's anyone on paid Datanika yet — this is the policy for when there is.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this doesn't mean
&lt;/h2&gt;

&lt;p&gt;We are not becoming a per-row-pricing company. We are not adding event fees. We are not going to count deletes. "Processed GB" is the only new meter. Everything else on your bill — seats, connections, schedules, support — stays on the subscription.&lt;/p&gt;

&lt;p&gt;If you've been evaluating Datanika on the v1 pricing page and waiting to decide: the economics on your bill are now predictable against your data volume, which is probably the thing you actually wanted.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers, measured
&lt;/h2&gt;

&lt;p&gt;On a standard Hetzner CPX32 (4 vCPU, 8 GB RAM, €13/mo), dlt — the library a Datanika upload runs — processed &lt;strong&gt;17,704 rows/second&lt;/strong&gt; on a 10.1M-row Postgres → DuckDB pipeline, full extract, normalize and load, with a p95 of 571s across 3 runs (569.9s, 570.5s, 571.4s). The benchmark script calls dlt directly rather than going through a Datanika upload, so it measures the library, not the product. Full benchmark log and methodology in &lt;a href="https://datanika.io/blog/datanika-vs-modern-data-stack/" rel="noopener noreferrer"&gt;Datanika vs. the Modern Data Stack&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://datanika.io/why-cheaper/" rel="noopener noreferrer"&gt;/why-cheaper/&lt;/a&gt; calculator lets you drag a slider from 1 GB to 10 TB and see the cost side-by-side with Fivetran Starter, with Datanika auto-picking the cheapest tier: Free up to 10 GB, then Pro-with-overage or Enterprise-flat, whichever costs less.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Try it free at &lt;a href="https://app.datanika.io" rel="noopener noreferrer"&gt;app.datanika.io&lt;/a&gt;&lt;/strong&gt; — 10 GB/mo on Free, no credit card, and one published per-GB rate with no MAR arithmetic to reverse-engineer.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Correction, 2026-08-31.&lt;/em&gt; An earlier version of this post said Pro and Enterprise pipelines show a pre-run &lt;code&gt;predicted_bytes&lt;/code&gt; estimate, derived from a moving average of the last five runs, before you click "Run." That was written from the pricing spec rather than from the product: the estimate is not computed today, on any plan. The paragraph has been replaced with what the meter actually does. Tracked in &lt;a href="https://github.com/datanika-io/datanika-landing/issues/375" rel="noopener noreferrer"&gt;landing#375&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Correction, 2026-09-22.&lt;/em&gt; Three more passages described something other than the product. The post had a section titled "Pick ELT, pay less". That section said every pipeline has an ETL/ELT mode selector, and that ELT is metered at ~0.8 GB where ETL is metered at ~3 GB. No pipeline can be switched to ELT: the app shows no mode selector, and nothing saves a mode, so every run takes the normalizing path described above. The worked example also counted a dbt model's 100 MB against your quota. It does not: only uploads are metered in bytes. And the benchmark was credited to Datanika when it measured dlt called directly by a script. All three passages now say what the product does. Separately, the post said Fivetran charges $5 per connection on top of MAR; Fivetran's own pricing page applies that base charge only to standard connections using between 1 and 1M MAR a month, and not on its Free plan, so the sentence now says so. Tracked in &lt;a href="https://github.com/datanika-io/datanika-landing/issues/656" rel="noopener noreferrer"&gt;landing#656&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>pricing</category>
      <category>uniteconomics</category>
      <category>volumepricing</category>
      <category>elt</category>
    </item>
    <item>
      <title>How We Built a 5.8x Faster Data Pipeline with DLT + Arrow</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Wed, 23 Sep 2026 12:35:11 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/how-we-built-a-58x-faster-data-pipeline-with-dlt-arrow-5el</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/how-we-built-a-58x-faster-data-pipeline-with-dlt-arrow-5el</guid>
      <description>&lt;p&gt;Before Datanika existed, the founding team ran production data pipelines at &lt;a href="https://whisk.com" rel="noopener noreferrer"&gt;Whisk&lt;/a&gt; — a food-tech platform processing 54 million MySQL rows and 8.6 million MongoDB documents daily into ClickHouse. We went through three generations of architecture, and the numbers tell the story better than any pitch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three generations
&lt;/h2&gt;

&lt;h3&gt;
  
  
  v1: Custom Airflow operators (99 min for 54M rows)
&lt;/h3&gt;

&lt;p&gt;Each source type had its own operator. &lt;code&gt;MySqlToClickHouseTransfer&lt;/code&gt;, &lt;code&gt;PostgresToClickHouseOperator&lt;/code&gt;, &lt;code&gt;MongoToCH&lt;/code&gt; — four files with duplicated logic, no shared schema management, no execution logging, no incremental state.&lt;/p&gt;

&lt;p&gt;It worked. For small tables it was fine. But the main table — &lt;code&gt;user_recipe_rel&lt;/code&gt;, 54 million rows — took 99 minutes as a batch INSERT over the native ClickHouse protocol. Acceptable, not fast.&lt;/p&gt;

&lt;h3&gt;
  
  
  v2: DLT + JSON normalization (174 min — slower)
&lt;/h3&gt;

&lt;p&gt;We unified everything behind DLT's &lt;code&gt;sql_database&lt;/code&gt; source and a single &lt;code&gt;DltRunOperator&lt;/code&gt;. YAML-driven table configs, automatic schema evolution, centralized execution logging, incremental state tracking. The operator improvement was night and day.&lt;/p&gt;

&lt;p&gt;The performance was not.&lt;/p&gt;

&lt;p&gt;DLT's default mode processes rows as Python dicts: fetch from source, convert each row to a dict, run schema inference, write to JSONL, convert to Parquet, upload. For 54 million rows, this JSON normalization pipeline took &lt;strong&gt;174 minutes&lt;/strong&gt; — 75% slower than the legacy batch INSERT.&lt;/p&gt;

&lt;p&gt;The operator was better. The throughput was worse.&lt;/p&gt;

&lt;h3&gt;
  
  
  v3: DLT + Arrow mode (17 min — 5.8x faster)
&lt;/h3&gt;

&lt;p&gt;Arrow mode changes the data flow completely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MySQL --&amp;gt; Server-side cursor --&amp;gt; PyArrow Tables --&amp;gt; Parquet Files --&amp;gt; ClickHouse
          (chunk_size rows)     (in memory)        (file_max_bytes)   (HTTP insert)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No dict conversion, no schema inference, no JSONL intermediate step. The source yields &lt;code&gt;pa.Table&lt;/code&gt; objects directly, DLT writes them to Parquet files, and the Parquet files upload to ClickHouse via HTTP.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;54 million rows in 17 minutes. 5.8x faster than legacy, 10x faster than DLT + JSON.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;h3&gt;
  
  
  MySQL: &lt;code&gt;user_recipe_rel&lt;/code&gt; (54M rows, production)
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Legacy operator&lt;/th&gt;
&lt;th&gt;DLT + JSON&lt;/th&gt;
&lt;th&gt;DLT + Arrow&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fetch from MySQL&lt;/td&gt;
&lt;td&gt;~99 min (bundled)&lt;/td&gt;
&lt;td&gt;~72 min&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~10 min&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Normalize + Load&lt;/td&gt;
&lt;td&gt;bundled&lt;/td&gt;
&lt;td&gt;~102 min&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~7 min&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total time&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;99 min&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;174 min&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;17 min&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Load files&lt;/td&gt;
&lt;td&gt;batch inserts&lt;/td&gt;
&lt;td&gt;547&lt;/td&gt;
&lt;td&gt;5,462&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Speedup vs legacy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;baseline&lt;/td&gt;
&lt;td&gt;0.57x (slower)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;5.8x faster&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Arrow fetch is 7x faster because it reads columnar batches instead of iterating row-by-row through Python dicts. The normalize step drops from 102 minutes to effectively zero because Arrow's normalizer does a direct Parquet file import — a file rename, not data processing.&lt;/p&gt;

&lt;h3&gt;
  
  
  MongoDB: &lt;code&gt;hostedImages&lt;/code&gt; (8.6M documents, dev)
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Legacy operator&lt;/th&gt;
&lt;th&gt;DLT + JSON&lt;/th&gt;
&lt;th&gt;DLT + Arrow&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fetch from MongoDB&lt;/td&gt;
&lt;td&gt;~49 min (prod, 29.7M)&lt;/td&gt;
&lt;td&gt;~11 min&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~3.3 min&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Normalize&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;td&gt;~12 min (JSONL)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~1 sec&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Load to ClickHouse&lt;/td&gt;
&lt;td&gt;bundled&lt;/td&gt;
&lt;td&gt;~1 min&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~40 sec&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;~49 min&lt;/strong&gt; (29.7M)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;~23.5 min&lt;/strong&gt; (8.6M)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;~4 min&lt;/strong&gt; (8.6M)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Normalized: sec/M rows&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;99s&lt;/td&gt;
&lt;td&gt;164s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;28s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;vs legacy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;baseline&lt;/td&gt;
&lt;td&gt;0.6x (slower)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.5x faster&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Extrapolated to production scale (29.7M documents), DLT + Arrow would complete in approximately 14 minutes versus the legacy operator's 49 minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The &lt;code&gt;file_max_bytes&lt;/code&gt; war story
&lt;/h2&gt;

&lt;p&gt;Arrow mode has one critical configuration that doesn't exist in JSON mode: &lt;code&gt;DATA_WRITER__FILE_MAX_BYTES&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Without it, DLT concatenates all Arrow chunks from the extract step into a single Parquet file. For 54 million rows, that's a multi-gigabyte file. And when you try to upload a 2 GB Parquet file via HTTP to ClickHouse through an nginx proxy, you get a timeout.&lt;/p&gt;

&lt;p&gt;The fix: set &lt;code&gt;DATA_WRITER__FILE_MAX_BYTES=100MB&lt;/code&gt;. This splits the extract output into manageable ~100 MB Parquet files — 5,462 files for the 54M-row table. Each uploads in seconds.&lt;/p&gt;

&lt;p&gt;The gotcha: the environment variable must use the &lt;code&gt;DATA_WRITER__&lt;/code&gt; prefix, not &lt;code&gt;NORMALIZE__DATA_WRITER__&lt;/code&gt;. Arrow's direct-import path skips the normalize step entirely, so normalize-scoped config has no effect. We discovered this after a week of "why is &lt;code&gt;NORMALIZE__DATA_WRITER__FILE_MAX_ITEMS&lt;/code&gt; not splitting files" debugging.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters for Datanika
&lt;/h2&gt;

&lt;p&gt;Datanika runs dlt, so the trade-off this post measured sits under every upload.&lt;/p&gt;

&lt;p&gt;There is no ELT mode to switch a Datanika pipeline to. Every upload runs through dlt's normalize step on our side, and the &lt;a href="https://datanika.io/features/volume-pricing/" rel="noopener noreferrer"&gt;meter&lt;/a&gt; counts what that step writes: a 1 GB JSON export that becomes ~3 GB of flat tables is metered as ~3 GB.&lt;/p&gt;

&lt;p&gt;Our &lt;a href="https://datanika.io/blog/datanika-vs-modern-data-stack/" rel="noopener noreferrer"&gt;CPX32 benchmark&lt;/a&gt; (10.1M rows, Postgres to DuckDB) is not an Arrow number either. Its script calls dlt with the default backend, not Arrow, and completes the full extract, normalize and load in 571 seconds at 17,704 rows/s. It measures dlt called directly, not a Datanika upload.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;DLT's JSON normalization is the bottleneck, not the source fetch.&lt;/strong&gt; Switching to Arrow skips it entirely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Arrow mode requires explicit file splitting.&lt;/strong&gt; Without &lt;code&gt;DATA_WRITER__FILE_MAX_BYTES&lt;/code&gt;, you'll build multi-GB files that timeout on upload.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The speedup is not from Arrow being faster at reading&lt;/strong&gt; — it's from eliminating the dict-to-JSON-to-Parquet conversion chain.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you're running dlt at scale and haven't tried Arrow mode yet, the performance gain is significant and the migration is straightforward — set &lt;code&gt;use_arrow: True&lt;/code&gt; in your source config and add the &lt;code&gt;DATA_WRITER__FILE_MAX_BYTES&lt;/code&gt; env var.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Datanika is an open-source data pipeline platform built on dlt + dbt-core. &lt;a href="https://app.datanika.io/" rel="noopener noreferrer"&gt;Start free&lt;/a&gt; with 10 GB/mo, &lt;a href="https://datanika.io/docs/self-hosting/" rel="noopener noreferrer"&gt;self-host it&lt;/a&gt;, or &lt;a href="https://datanika.io/features/volume-pricing/" rel="noopener noreferrer"&gt;read about our volume-based pricing&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Correction, 2026-09-22.&lt;/em&gt; An earlier version of this post said this architecture powers a Datanika "ELT mode" you can select on a pipeline, metered at ~0.8 GB where ETL is metered at ~3 GB. There is no such mode to select: the app shows no mode selector and nothing saves a mode, so every upload runs the normalizing path. The post also said the CPX32 benchmark was built on this Arrow path, but that script uses dlt's default backend. The "Why this matters for Datanika" section has been rewritten to say what the product does, and a fourth takeaway that repeated the claim is gone. Tracked in &lt;a href="https://github.com/datanika-io/datanika-landing/issues/656" rel="noopener noreferrer"&gt;landing#656&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>dlt</category>
      <category>arrow</category>
      <category>performance</category>
      <category>benchmark</category>
    </item>
    <item>
      <title>How to Build a Stripe Revenue Dashboard with dbt</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Tue, 22 Sep 2026 14:38:59 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/how-to-build-a-stripe-revenue-dashboard-with-dbt-40ek</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/how-to-build-a-stripe-revenue-dashboard-with-dbt-40ek</guid>
      <description>&lt;p&gt;Stripe's built-in dashboard is genuinely good — for operations. You can see today's charges, chase a failed payment, and eyeball gross volume. But the moment finance asks "what was net MRR in March, split by plan, excluding trials and after refunds?" you're exporting CSVs and fighting a pivot table. Stripe's reporting is deliberately shallow because Stripe is a payments processor, not an analytics warehouse.&lt;/p&gt;

&lt;p&gt;The fix is the same one every serious SaaS team eventually lands on: &lt;strong&gt;get Stripe's raw objects into a warehouse, model them with dbt, and point a BI tool at the result.&lt;/strong&gt; Once the data is modeled, MRR, churn, ARR, LTV, and cohort retention are just SQL — recomputed every morning, joinable against product usage, and auditable line by line.&lt;/p&gt;

&lt;p&gt;This is the end-to-end tutorial. We'll land Stripe in a warehouse with &lt;a href="https://datanika.io/connectors/stripe/" rel="noopener noreferrer"&gt;Datanika&lt;/a&gt; (which wraps &lt;code&gt;dlt&lt;/code&gt; for extract and &lt;code&gt;dbt-core&lt;/code&gt; for transform in one app), write the staging and mart models, add tests, schedule the refresh, and connect a dashboard. Every SQL snippet below is real and runnable — adjust the column names to match what actually lands in your warehouse and you're done.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we're building
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Stripe API
   │  (dlt, via Datanika)
   ▼
raw_stripe.*          ← raw objects: customers, invoices, subscriptions, charges
   │  (dbt staging: clean + cast)
   ▼
stg_stripe__*         ← typed, renamed, cents → dollars, epoch → timestamp
   │  (dbt marts: business logic)
   ▼
revenue_by_month      ┐
active_subscriptions  ├→  Metabase / Looker / your BI tool
customer_ltv          ┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three layers, one principle: &lt;strong&gt;raw data lands untouched, staging cleans it, marts hold the business logic.&lt;/strong&gt; When someone disputes an MRR number six months from now, you can trace it from the dashboard tile all the way back to a specific Stripe invoice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1 — Land Stripe in your warehouse
&lt;/h2&gt;

&lt;p&gt;We won't repeat the full connector walkthrough here — the &lt;a href="https://datanika.io/docs/connectors/stripe/" rel="noopener noreferrer"&gt;Stripe setup guide&lt;/a&gt; covers it end to end (create a read-only restricted key, add the connection, pick resources, run, schedule). The short version:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;In Stripe, create a &lt;strong&gt;restricted key&lt;/strong&gt; (&lt;code&gt;Developers → API keys → Create restricted key&lt;/code&gt;) with &lt;strong&gt;Read&lt;/strong&gt; permission on &lt;code&gt;Customers&lt;/code&gt;, &lt;code&gt;Charges&lt;/code&gt;, &lt;code&gt;Invoices&lt;/code&gt;, &lt;code&gt;Subscriptions&lt;/code&gt;, &lt;code&gt;Products&lt;/code&gt;, and &lt;code&gt;Prices&lt;/code&gt;. Never use a standard secret key — Datanika only ever reads.&lt;/li&gt;
&lt;li&gt;In Datanika, open &lt;strong&gt;&lt;code&gt;/connections&lt;/code&gt;&lt;/strong&gt;, pick &lt;strong&gt;Stripe&lt;/strong&gt;, and paste the key (stored encrypted at rest with Fernet).&lt;/li&gt;
&lt;li&gt;Create the upload. Because Stripe is a SaaS source, the form shows &lt;strong&gt;Select endpoints to load&lt;/strong&gt; — a checkbox each for &lt;code&gt;charges&lt;/code&gt;, &lt;code&gt;customers&lt;/code&gt;, &lt;code&gt;invoices&lt;/code&gt;, &lt;code&gt;prices&lt;/code&gt;, &lt;code&gt;products&lt;/code&gt; and &lt;code&gt;subscriptions&lt;/code&gt;, all ticked by default. Untick what you don't need; each ticked endpoint becomes its own table.&lt;/li&gt;
&lt;li&gt;Run it from the &lt;strong&gt;&lt;code&gt;/uploads&lt;/code&gt;&lt;/strong&gt; row for your upload — the button is &lt;strong&gt;Run&lt;/strong&gt;, and the trigger lives on the upload's own row, not on a pipeline page. Watch it in &lt;strong&gt;Runs&lt;/strong&gt;; the &lt;strong&gt;Rows&lt;/strong&gt; column there is one total for the whole run, not a per-endpoint breakdown.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where it lands, and why there's no field for it.&lt;/strong&gt; A SaaS source has &lt;strong&gt;no write-disposition, load-mode, source-schema or table-name control&lt;/strong&gt; — those are rendered only for SQL database sources, and the endpoint checkboxes are the equivalent. So the destination schema is derived, not typed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;BigQuery, Snowflake, Databricks&lt;/strong&gt; — it's the &lt;strong&gt;Dataset&lt;/strong&gt; / &lt;strong&gt;Schema&lt;/strong&gt; field on the &lt;em&gt;destination connection&lt;/em&gt;, not on the upload.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Postgres, MySQL, DuckDB and the rest&lt;/strong&gt; — it's the &lt;strong&gt;upload's own name, snake-cased&lt;/strong&gt;. Upload names accept letters, digits and spaces only, so you can't type &lt;code&gt;raw_stripe&lt;/code&gt; directly: name the upload &lt;strong&gt;&lt;code&gt;Raw Stripe&lt;/code&gt;&lt;/strong&gt; and you get the schema &lt;code&gt;raw_stripe&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This tutorial assumes you did one of those two things and ended up with &lt;code&gt;raw_stripe&lt;/code&gt;. If you named the upload something else, substitute your schema everywhere below.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When it finishes, open &lt;strong&gt;Models&lt;/strong&gt; — the sidebar entry for Datanika's data catalog, at &lt;strong&gt;&lt;code&gt;/models&lt;/code&gt;&lt;/strong&gt;. You should see one row per landed table: &lt;code&gt;customers&lt;/code&gt;, &lt;code&gt;invoices&lt;/code&gt;, &lt;code&gt;subscriptions&lt;/code&gt;, &lt;code&gt;charges&lt;/code&gt;, and so on, each with its &lt;strong&gt;Schema&lt;/strong&gt; and &lt;strong&gt;Columns&lt;/strong&gt;. dlt's &lt;code&gt;_dlt_loads&lt;/code&gt; / &lt;code&gt;_dlt_pipeline_state&lt;/code&gt; / &lt;code&gt;_dlt_version&lt;/code&gt; bookkeeping tables are created in the warehouse but filtered out of the catalog, so their absence here is expected, not a failed run. &lt;strong&gt;Confirm the schema name in the Schema column before you write &lt;code&gt;sources.yml&lt;/code&gt;&lt;/strong&gt; — it is the single most common reason the models below fail to resolve. Keep that tab open; you'll use it to check exact column names as you write models too.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Two things to know about Stripe's data before you model it.&lt;/strong&gt; (1) &lt;strong&gt;Amounts are integers in the smallest currency unit&lt;/strong&gt; — &lt;code&gt;2000&lt;/code&gt; means $20.00 in USD. You divide by 100 in staging (except zero-decimal currencies like JPY — more on that below). (2) &lt;strong&gt;Timestamps are Unix epoch seconds&lt;/strong&gt; — &lt;code&gt;1719792000&lt;/code&gt;, not &lt;code&gt;2024-07-01&lt;/code&gt;. &lt;code&gt;dlt&lt;/code&gt; sometimes types these as proper timestamps during normalization and sometimes lands them as integers; check the column's type in &lt;strong&gt;Models&lt;/strong&gt; and convert in staging if needed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Step 2 — Declare the source
&lt;/h2&gt;

&lt;p&gt;In Datanika, dbt models live under &lt;strong&gt;Transformations&lt;/strong&gt;. Start by telling dbt where the raw data is. Create a &lt;code&gt;sources.yml&lt;/code&gt; (or add to your existing one):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# models/staging/stripe/sources.yml&lt;/span&gt;
&lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;

&lt;span class="na"&gt;sources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;raw_stripe&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Raw&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Stripe&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;objects&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;loaded&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;by&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Datanika/dlt"&lt;/span&gt;
    &lt;span class="na"&gt;tables&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;customers&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;invoices&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;subscriptions&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;charges&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;prices&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now &lt;code&gt;{{ source('raw_stripe', 'invoices') }}&lt;/code&gt; resolves to the landed table, and dbt tracks the dependency in its DAG.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3 — Staging models: clean and cast
&lt;/h2&gt;

&lt;p&gt;Staging models are thin views, one per raw table. They do exactly four things: &lt;strong&gt;rename&lt;/strong&gt; to your conventions, &lt;strong&gt;cast&lt;/strong&gt; types, &lt;strong&gt;convert&lt;/strong&gt; cents and epochs, and &lt;strong&gt;drop&lt;/strong&gt; columns you'll never use. No business logic — that's the marts' job.&lt;/p&gt;

&lt;p&gt;Stripe stores foreign keys under the object's own name (&lt;code&gt;customer&lt;/code&gt;, &lt;code&gt;subscription&lt;/code&gt;), so we rename them to &lt;code&gt;_id&lt;/code&gt; for clarity.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/staging/stripe/stg_stripe__invoices.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'view'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt;                                          &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;invoice_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;customer&lt;/span&gt;                                    &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;subscription&lt;/span&gt;                                &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;subscription_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;currency&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="c1"&gt;-- cents → major units. Guard zero-decimal currencies (JPY, KRW…):&lt;/span&gt;
    &lt;span class="n"&gt;amount_paid&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;                         &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;amount_paid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;amount_due&lt;/span&gt;  &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;                         &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;amount_due&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;total&lt;/span&gt;       &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;                         &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="c1"&gt;-- Unix epoch → timestamp (BigQuery). See dialect notes below.&lt;/span&gt;
    &lt;span class="n"&gt;timestamp_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;created&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                  &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timestamp_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;period_start&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;             &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;period_start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timestamp_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;period_end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;               &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;period_end&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'raw_stripe'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'invoices'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/staging/stripe/stg_stripe__subscriptions.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'view'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt;                                          &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;subscription_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;customer&lt;/span&gt;                                    &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                                     &lt;span class="c1"&gt;-- active, trialing, past_due, canceled…&lt;/span&gt;
    &lt;span class="n"&gt;timestamp_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;created&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                  &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timestamp_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current_period_start&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;current_period_start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timestamp_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current_period_end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;       &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;current_period_end&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="k"&gt;when&lt;/span&gt; &lt;span class="n"&gt;canceled_at&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;
         &lt;span class="k"&gt;then&lt;/span&gt; &lt;span class="n"&gt;timestamp_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;canceled_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;canceled_at&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'raw_stripe'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'subscriptions'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/staging/stripe/stg_stripe__customers.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'view'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt;                          &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;currency&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timestamp_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;created&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'raw_stripe'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'customers'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warehouse dialect.&lt;/strong&gt; The &lt;code&gt;timestamp_seconds()&lt;/code&gt; above is BigQuery. On &lt;strong&gt;Postgres/Redshift&lt;/strong&gt; use &lt;code&gt;to_timestamp(created)&lt;/code&gt;; on &lt;strong&gt;Snowflake&lt;/strong&gt; use &lt;code&gt;to_timestamp(created)&lt;/code&gt;; on &lt;strong&gt;DuckDB&lt;/strong&gt; use &lt;code&gt;to_timestamp(created)&lt;/code&gt; or &lt;code&gt;make_timestamp(created * 1000000)&lt;/code&gt;. If &lt;code&gt;dlt&lt;/code&gt; already landed these as timestamps, drop the wrapper entirely. This is the one place warehouse portability actually bites — everything downstream is plain SQL.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Step 4 — The revenue marts
&lt;/h2&gt;

&lt;p&gt;Now the fun part. Each mart is a &lt;code&gt;table&lt;/code&gt; (materialized, because BI tools query it repeatedly).&lt;/p&gt;

&lt;h3&gt;
  
  
  Collected revenue by month
&lt;/h3&gt;

&lt;p&gt;The most-asked-for chart: how much did we actually collect, by month? Paid invoices are the source of truth — a charge can succeed without an invoice, and an invoice can exist without being paid, so we filter to &lt;code&gt;status = 'paid'&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/marts/finance/revenue_by_month.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'table'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;timestamp_trunc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;period_start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;-- BigQuery TIMESTAMP_TRUNC; Postgres: date_trunc('month', period_start)&lt;/span&gt;
    &lt;span class="n"&gt;currency&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;distinct&lt;/span&gt; &lt;span class="n"&gt;invoice_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;       &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;invoices_paid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;distinct&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;paying_customers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount_paid&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                 &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;revenue&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_stripe__invoices'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'paid'&lt;/span&gt;
&lt;span class="k"&gt;group&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
&lt;span class="k"&gt;order&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's recognized, collected revenue — the number finance reconciles against the bank. Group by &lt;code&gt;period_start&lt;/code&gt; (the service period) rather than &lt;code&gt;created_at&lt;/code&gt; if you want revenue attributed to the month it covers instead of the month it was billed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Active subscriptions and churn
&lt;/h3&gt;

&lt;p&gt;MRR movement starts with counting subscriptions in each state. This model gives you active counts and a monthly churn signal straight from the top-level subscription fields:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/marts/finance/subscription_movement.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'table'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;subs&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_stripe__subscriptions'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;

&lt;span class="n"&gt;started&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;timestamp_trunc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
           &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;new_subscriptions&lt;/span&gt;
    &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;subs&lt;/span&gt;
    &lt;span class="k"&gt;group&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;

&lt;span class="n"&gt;churned&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;timestamp_trunc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;canceled_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
           &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;churned_subscriptions&lt;/span&gt;
    &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;subs&lt;/span&gt;
    &lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;canceled_at&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;
    &lt;span class="k"&gt;group&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="n"&gt;coalesce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                        &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;coalesce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;new_subscriptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                  &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;new_subscriptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;coalesce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;churned_subscriptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;              &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;churned_subscriptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;coalesce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;new_subscriptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;coalesce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;churned_subscriptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;        &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;net_new_subscriptions&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;started&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
&lt;span class="k"&gt;full&lt;/span&gt; &lt;span class="k"&gt;outer&lt;/span&gt; &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="n"&gt;churned&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;order&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="k"&gt;month&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a point-in-time active count, filter &lt;code&gt;stg_stripe__subscriptions&lt;/code&gt; to &lt;code&gt;status in ('active', 'trialing')&lt;/code&gt; — that's your live subscriber base right now.&lt;/p&gt;

&lt;h3&gt;
  
  
  Customer lifetime value
&lt;/h3&gt;

&lt;p&gt;LTV to date is just total collected per customer. Join to &lt;code&gt;customers&lt;/code&gt; for the email/name so the dashboard is readable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/marts/finance/customer_ltv.sql&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'table'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;created_at&lt;/span&gt;                     &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;signed_up_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;distinct&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;invoice_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;invoices_paid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;amount_paid&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;               &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;lifetime_revenue&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;first_payment_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;last_payment_at&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_stripe__customers'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
&lt;span class="k"&gt;left&lt;/span&gt; &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_stripe__invoices'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;
       &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;
      &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'paid'&lt;/span&gt;
&lt;span class="k"&gt;group&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
&lt;span class="k"&gt;order&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="n"&gt;lifetime_revenue&lt;/span&gt; &lt;span class="k"&gt;desc&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  True MRR from subscription items (the honest caveat)
&lt;/h3&gt;

&lt;p&gt;Everything above uses top-level, high-confidence columns. &lt;strong&gt;Normalized MRR is the one metric where your schema will differ from mine&lt;/strong&gt;, because a subscription's price lives in a &lt;em&gt;nested&lt;/em&gt; array (&lt;code&gt;items.data[]&lt;/code&gt;), and &lt;code&gt;dlt&lt;/code&gt; flattens nested arrays into a child table — often something like &lt;code&gt;subscriptions__items__data&lt;/code&gt;. Open &lt;strong&gt;Models&lt;/strong&gt; and find the real name before writing this model.&lt;/p&gt;

&lt;p&gt;Once you've located it, MRR is: for every active subscription item, take &lt;code&gt;unit_amount × quantity&lt;/code&gt;, normalize it to a monthly figure by the price's billing interval, and sum:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- models/marts/finance/mrr.sql  ── adjust the source table name to the one in Models&lt;/span&gt;
&lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;materialized&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'table'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;unit_amount&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;quantity&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;recurring_interval&lt;/span&gt;          &lt;span class="c1"&gt;-- price's recurring.interval&lt;/span&gt;
              &lt;span class="k"&gt;when&lt;/span&gt; &lt;span class="s1"&gt;'month'&lt;/span&gt; &lt;span class="k"&gt;then&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;
              &lt;span class="k"&gt;when&lt;/span&gt; &lt;span class="s1"&gt;'year'&lt;/span&gt;  &lt;span class="k"&gt;then&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;12&lt;/span&gt;
              &lt;span class="k"&gt;when&lt;/span&gt; &lt;span class="s1"&gt;'week'&lt;/span&gt;  &lt;span class="k"&gt;then&lt;/span&gt; &lt;span class="mi"&gt;52&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;12&lt;/span&gt;
              &lt;span class="k"&gt;when&lt;/span&gt; &lt;span class="s1"&gt;'day'&lt;/span&gt;   &lt;span class="k"&gt;then&lt;/span&gt; &lt;span class="mi"&gt;365&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;12&lt;/span&gt;
          &lt;span class="k"&gt;end&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;mrr&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'raw_stripe'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'subscription_items'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;   &lt;span class="c1"&gt;-- ← your actual child-table name&lt;/span&gt;
&lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'stg_stripe__subscriptions'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
     &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;subscription_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;subscription_id&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'active'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'trialing'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If most of your subscriptions carry a single price (typical for simpler SaaS), you can skip the nested table and read the price fields directly off the subscription. Either way, the interval-normalization &lt;code&gt;CASE&lt;/code&gt; is the part people get wrong — a yearly plan is not $1,200 of MRR, it's $100.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5 — Test what you built
&lt;/h2&gt;

&lt;p&gt;dbt tests turn "I think this is right" into CI. Add a &lt;code&gt;schema.yml&lt;/code&gt; next to your models:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# models/marts/finance/schema.yml&lt;/span&gt;
&lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;

&lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;revenue_by_month&lt;/span&gt;
    &lt;span class="na"&gt;columns&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;revenue&lt;/span&gt;
        &lt;span class="na"&gt;tests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;not_null&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;customer_ltv&lt;/span&gt;
    &lt;span class="na"&gt;columns&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;customer_id&lt;/span&gt;
        &lt;span class="na"&gt;tests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;not_null&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;unique&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run them from the Transformations UI. A failing &lt;code&gt;unique&lt;/code&gt; on &lt;code&gt;customer_id&lt;/code&gt; means a fan-out bug in your join; a &lt;code&gt;not_null&lt;/code&gt; on &lt;code&gt;revenue&lt;/code&gt; means an invoice slipped through with a null amount. Catch it here, not in a board deck.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6 — Schedule the whole pipeline
&lt;/h2&gt;

&lt;p&gt;A dashboard is only as fresh as its slowest step. In Datanika, wire the extract and the transform into one dependency chain so transforms run &lt;strong&gt;only after&lt;/strong&gt; the Stripe load succeeds:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Schedule the &lt;strong&gt;Stripe upload&lt;/strong&gt; — hourly for revenue-ops dashboards, or daily at &lt;code&gt;03:00&lt;/code&gt; if Stripe is one of many batch sources. (See the &lt;a href="https://datanika.io/docs/scheduling/" rel="noopener noreferrer"&gt;Scheduling guide&lt;/a&gt; for cron syntax and timezones.)&lt;/li&gt;
&lt;li&gt;Schedule the &lt;strong&gt;transformation pipeline&lt;/strong&gt; to depend on that upload. Datanika's DAG guarantees dbt won't run against half-loaded data.&lt;/li&gt;
&lt;li&gt;Wire failure alerts in &lt;strong&gt;Settings → Notifications&lt;/strong&gt; so a broken run pages your team before finance notices the numbers are stale — here's the &lt;a href="https://datanika.io/blog/slack-alerts-pipeline-failures/" rel="noopener noreferrer"&gt;Slack alerts walkthrough&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Step 7 — Point a dashboard at the marts
&lt;/h2&gt;

&lt;p&gt;Your marts are just tables in the warehouse now, so any BI tool works. With &lt;strong&gt;Metabase&lt;/strong&gt; (open source, free, runs on the same box if you self-host):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Revenue trend&lt;/strong&gt; — line chart on &lt;code&gt;revenue_by_month&lt;/code&gt; (&lt;code&gt;month&lt;/code&gt; × &lt;code&gt;revenue&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MRR&lt;/strong&gt; — a single big-number tile on &lt;code&gt;mrr&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Net subscriber movement&lt;/strong&gt; — bar chart on &lt;code&gt;subscription_movement&lt;/code&gt; (&lt;code&gt;new&lt;/code&gt; vs &lt;code&gt;churned&lt;/code&gt; per month).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Top customers&lt;/strong&gt; — table on &lt;code&gt;customer_ltv&lt;/code&gt; sorted by &lt;code&gt;lifetime_revenue&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because the logic lives in dbt, not in Metabase, every tile shares one definition of "revenue." Swap Metabase for Looker or Superset tomorrow and the numbers don't move.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you get
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Metrics Stripe won't give you&lt;/strong&gt; — MRR, churn, net revenue retention, LTV, cohorts — defined once in SQL and recomputed every morning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A full audit trail&lt;/strong&gt; — dashboard tile → mart → staging → the exact Stripe invoice. No black-box reporting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Joinable data&lt;/strong&gt; — Stripe now sits in the same warehouse as your product usage and CRM, so "revenue by feature adoption" becomes one query.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No per-seat transform tax&lt;/strong&gt; — it's &lt;code&gt;dbt-core&lt;/code&gt;, running on your schedule.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;Datanika meters &lt;strong&gt;bytes processed&lt;/strong&gt;, not rows or connectors. The &lt;a href="https://datanika.io/pricing/" rel="noopener noreferrer"&gt;Free plan&lt;/a&gt; includes &lt;strong&gt;10 GB/month&lt;/strong&gt;. Pro is &lt;strong&gt;$79/mo with 100 GB included&lt;/strong&gt; and &lt;strong&gt;$0.50/GB&lt;/strong&gt; beyond that. A Stripe account's daily refresh is small — Stripe objects are narrow JSON — so for most teams the bill here is the subscription, not the volume. Don't take that on trust, and don't look for the answer in the app: its &lt;strong&gt;Plan Usage&lt;/strong&gt; panel counts model runs, not bytes. Size it from your data instead — the meter counts what an upload writes after normalization.&lt;/p&gt;

&lt;p&gt;Compare that to &lt;a href="https://datanika.io/compare/fivetran/" rel="noopener noreferrer"&gt;Fivetran&lt;/a&gt;, which counts each Stripe object as monthly active rows and adds a per-connection minimum — the exact pricing model this whole tutorial routes around.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next steps
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/docs/connectors/stripe/" rel="noopener noreferrer"&gt;Connect Stripe →&lt;/a&gt;&lt;/strong&gt; the full setup guide (restricted key, resources, first run).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/use-cases/stripe-to-bigquery/" rel="noopener noreferrer"&gt;Stripe → BigQuery use case&lt;/a&gt;&lt;/strong&gt; — the reference architecture for this pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/docs/transformations/" rel="noopener noreferrer"&gt;Transformations guide&lt;/a&gt;&lt;/strong&gt; — models, tests, snapshots, and materializations in Datanika.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/connectors/" rel="noopener noreferrer"&gt;Browse all connectors&lt;/a&gt;&lt;/strong&gt; — join Stripe against HubSpot, your product Postgres, Facebook Ads, and more.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Stripe gives you the payments. dbt gives you the metrics. Datanika is the ten-minute bridge between them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://app.datanika.io/" rel="noopener noreferrer"&gt;Start free at app.datanika.io&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Correction, 2026-09-22.&lt;/em&gt; The cost section told you to check your byte usage in the dashboard's Plan Usage panel. That panel counts model runs, not bytes, and no screen in the app shows your byte count today, so the section now says to size it from your data. Tracked in &lt;a href="https://github.com/datanika-io/datanika-core/issues/1513" rel="noopener noreferrer"&gt;datanika-core#1513&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>tutorial</category>
      <category>dbt</category>
      <category>dataengineering</category>
      <category>stripe</category>
    </item>
    <item>
      <title>Your Data Pipelines, Now an MCP Server</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Mon, 21 Sep 2026 10:36:31 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/your-data-pipelines-now-an-mcp-server-1p29</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/your-data-pipelines-now-an-mcp-server-1p29</guid>
      <description>&lt;p&gt;Datanika now speaks MCP two ways. Paste our hosted endpoint into a client that supports remote MCP servers and approve it in the browser:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://app.datanika.io/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or run it locally — &lt;code&gt;datanika-mcp&lt;/code&gt; is on &lt;a href="https://pypi.org/project/datanika-mcp/" rel="noopener noreferrer"&gt;PyPI&lt;/a&gt; and listed on the &lt;a href="https://registry.modelcontextprotocol.io" rel="noopener noreferrer"&gt;official MCP registry&lt;/a&gt; as &lt;code&gt;io.datanika/datanika-mcp&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uvx datanika-mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read-only by default on both. Full setup is in the &lt;a href="https://datanika.io/docs/mcp-server/" rel="noopener noreferrer"&gt;MCP server guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What your agent can actually do
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://modelcontextprotocol.io" rel="noopener noreferrer"&gt;Model Context Protocol&lt;/a&gt; is how an LLM client gets typed tools instead of a blob of API documentation and a hope. We expose &lt;strong&gt;25 tools&lt;/strong&gt; over the &lt;a href="https://datanika.io/api/reference/" rel="noopener noreferrer"&gt;Datanika REST API&lt;/a&gt; — 17 read-only, plus 8 write tools that stay switched off until you ask for them.&lt;/p&gt;

&lt;p&gt;The interesting part isn't the count. It's that the read-only set is genuinely enough to be useful:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Discover and introspect&lt;/strong&gt; — list connections, read the typed config schema for any connector, list schemas and tables inside a source, preview rows, run a read-only &lt;code&gt;SELECT&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validate&lt;/strong&gt; — compile a dbt model, which resolves Jinja, &lt;code&gt;ref()&lt;/code&gt; and &lt;code&gt;source()&lt;/code&gt; without touching your warehouse. Then preview its output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observe&lt;/strong&gt; — list uploads, pipelines, transformations and runs; fetch a specific run and &lt;em&gt;its logs&lt;/em&gt;; browse the data catalog.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one changes the texture of debugging. "Why did last night's sync fail?" stops being a tab-switching expedition and becomes a question you ask in the same window where you're already working. The agent pulls the run, reads the logs, cross-references the connection config, and tells you the credential expired.&lt;/p&gt;

&lt;h2&gt;
  
  
  Read-only by default, and why that matters
&lt;/h2&gt;

&lt;p&gt;The 8 write tools — create a connection, create a model, trigger a run — refuse to execute unless you have deliberately enabled them. How deliberate depends on which path you took:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Over the hosted endpoint, write access is granted at authorization time.&lt;/strong&gt; When a client asks for write scopes, you approve them in the browser once, and the key that approval mints carries exactly those scopes. A client that asks for nothing gets read-only — silence is never read as consent to write — and a pasted API key stays read-only on that endpoint even if its own scopes would allow writes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Locally, writes are an explicit opt-in:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uvx datanika-mcp &lt;span class="nt"&gt;--allow-write&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Either way the default is read-only, and enabling writes is a decision someone has to make on purpose rather than a setting that drifts on.&lt;/p&gt;

&lt;p&gt;This is not security theatre, but it is also not the whole story, so here is the honest version. Authentication is the real ceiling: the server can't reach another organization or escalate its own scope, &lt;code&gt;query_connection&lt;/code&gt; is &lt;code&gt;SELECT&lt;/code&gt;-only enforced server-side, and everything an agent does lands in your &lt;a href="https://datanika.io/docs/audit-log/" rel="noopener noreferrer"&gt;audit log&lt;/a&gt;. The read-only default just means the &lt;em&gt;common&lt;/em&gt; case — "help me understand why this pipeline is broken" — needs no write access at all, so you shouldn't grant it.&lt;/p&gt;

&lt;p&gt;If you use the local path, use a dedicated API key for agent access rather than reusing one from CI. If an agent does something surprising, you revoke one key instead of untangling which automation broke. Over the hosted path there's no key to manage: consent mints one for you and revoking the grant is the off switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the agent leaves behind
&lt;/h2&gt;

&lt;p&gt;There's a reason we shipped this as tools over dbt and dlt rather than as a natural-language pipeline builder with its own format.&lt;/p&gt;

&lt;p&gt;An agent with write access can produce a great deal of work very quickly. If that work lands in a proprietary transformation format, the speed is a liability: you've automated the creation of something only one vendor can run. What &lt;code&gt;create_transformation&lt;/code&gt; writes here is a dbt model — SQL with &lt;code&gt;ref()&lt;/code&gt; and &lt;code&gt;source()&lt;/code&gt;, testable, versionable, and runnable by plain &lt;code&gt;dbt run&lt;/code&gt; on a machine that has never heard of Datanika.&lt;/p&gt;

&lt;p&gt;We wrote about that trade-off in more detail on the &lt;a href="https://datanika.io/ai-agents/#portability" rel="noopener noreferrer"&gt;AI agents page&lt;/a&gt;. It's the reason the MCP server is deliberately thin: it forwards to the REST API, which orchestrates dlt and dbt-core. There is no clever middle layer accumulating state you can't take with you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting it up
&lt;/h2&gt;

&lt;p&gt;If your client supports remote MCP servers, there is nothing to set up beyond pasting &lt;code&gt;https://app.datanika.io/mcp&lt;/code&gt; and approving the consent screen. Everything below is the local path.&lt;/p&gt;

&lt;p&gt;Claude Desktop, in &lt;code&gt;claude_desktop_config.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"datanika"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"uvx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"datanika-mcp"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--url"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://app.datanika.io"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--api-key"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"etf_your_key_here"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Claude Code is one command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add datanika &lt;span class="nt"&gt;--&lt;/span&gt; uvx datanika-mcp &lt;span class="nt"&gt;--api-key&lt;/span&gt; etf_your_key_here
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cursor reads &lt;code&gt;~/.cursor/mcp.json&lt;/code&gt; and takes the same shape, or environment variables if you prefer. Self-hosting? Point &lt;code&gt;--url&lt;/code&gt; at whatever origin serves &lt;code&gt;/api/v1&lt;/code&gt; — with the default Docker Compose setup that's &lt;code&gt;http://localhost:8000&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Get an API key from &lt;strong&gt;Settings → API Keys&lt;/strong&gt;, then see the &lt;a href="https://datanika.io/docs/mcp-server/" rel="noopener noreferrer"&gt;full guide&lt;/a&gt; for the per-client details, the complete tool table, and troubleshooting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this fits
&lt;/h2&gt;

&lt;p&gt;The MCP server joins the &lt;code&gt;/llms.txt&lt;/code&gt; discovery document and the &lt;a href="https://datanika.io/docs/ai-agents/" rel="noopener noreferrer"&gt;agent guide&lt;/a&gt; rather than replacing them. If you're driving Datanika from a script, a CI job, or your own agent framework, the REST API is still the direct path and the discovery documents still tell an LLM everything it needs. MCP is for the case where your agent already lives in a client that speaks it — which, increasingly, it does.&lt;/p&gt;

&lt;p&gt;The server is open source under AGPL-3.0, in the &lt;a href="https://github.com/datanika-io/datanika-core/tree/master/datanika-mcp" rel="noopener noreferrer"&gt;&lt;code&gt;datanika-mcp/&lt;/code&gt;&lt;/a&gt; directory of the core repo. Issues and PRs welcome.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Datanika is an open-source data pipeline platform built on dlt and dbt-core. &lt;a href="https://app.datanika.io/" rel="noopener noreferrer"&gt;Start free&lt;/a&gt; with 10 GB/month, or &lt;a href="https://datanika.io/docs/self-hosting/" rel="noopener noreferrer"&gt;self-host it&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>aiagents</category>
      <category>claude</category>
      <category>dbt</category>
    </item>
    <item>
      <title>Self-Hosting Datanika with Docker Compose</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Sun, 20 Sep 2026 21:17:18 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/self-hosting-datanika-with-docker-compose-2gob</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/self-hosting-datanika-with-docker-compose-2gob</guid>
      <description>&lt;p&gt;Most data platforms make self-hosting a footnote — a "contact sales for the on-prem option" link that goes nowhere, or an open-source core so stripped-down it's really just a demo for the paid cloud. Datanika's open-source core is the actual product: extract (&lt;code&gt;dlt&lt;/code&gt;), transform (&lt;code&gt;dbt-core&lt;/code&gt;), scheduling, a visual pipeline builder, every connector, multi-org RBAC, nine languages. It runs from one Docker Compose file, and self-hosting it costs &lt;strong&gt;$0 forever&lt;/strong&gt; — the AGPL-3.0 core has no license key, no seat count, and no GB meter.&lt;/p&gt;

&lt;p&gt;This is the walkthrough: what you get, the five-minute quick start, the configuration that actually matters, and — because this is the part most tutorials skip — an honest checklist for running it in production, where &lt;em&gt;you&lt;/em&gt; own the pager.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why self-host at all
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://app.datanika.io/" rel="noopener noreferrer"&gt;managed version at app.datanika.io&lt;/a&gt; exists because plenty of teams would rather not run infrastructure. Self-hosting is the right call when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Data residency / compliance.&lt;/strong&gt; Your customer data never leaves your VPC. No third-party processor, no data-processing addendum to negotiate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost at volume.&lt;/strong&gt; The managed plan meters bytes processed; self-hosted meters nothing. If you're moving terabytes a month, a $12 VPS beats any per-GB bill. (We did that math in &lt;a href="https://datanika.io/blog/real-cost-modern-data-stack/" rel="noopener noreferrer"&gt;The Real Cost of Your Modern Data Stack&lt;/a&gt;.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No vendor lock-in.&lt;/strong&gt; It's &lt;code&gt;dlt&lt;/code&gt; + &lt;code&gt;dbt-core&lt;/code&gt; under an open UI. If Datanika vanished tomorrow, your pipelines are standard dlt sources and dbt models — they keep running.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You already have infra.&lt;/strong&gt; A spare box, a Kubernetes cluster, a managed Postgres — drop Datanika next to them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tradeoff is real and we'll be straight about it below: self-hosting means you own upgrades, backups, and the 3 AM page when a disk fills up. For a lot of teams that's a fair trade. For some it isn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Docker Engine 24+&lt;/strong&gt; and &lt;strong&gt;Docker Compose v2&lt;/strong&gt; (&lt;code&gt;docker compose&lt;/code&gt;, not the old &lt;code&gt;docker-compose&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4 GB RAM&lt;/strong&gt; minimum, &lt;strong&gt;8 GB&lt;/strong&gt; recommended&lt;/li&gt;
&lt;li&gt;That's it. Postgres 16 and Redis 7 ship inside the Compose file — you don't install them separately unless you want to bring your own.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The five-minute quick start
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/datanika-io/datanika-core.git
&lt;span class="nb"&gt;cd &lt;/span&gt;datanika-core
&lt;span class="nb"&gt;cp&lt;/span&gt; .env.example .env
&lt;span class="c"&gt;# edit .env — at minimum set SECRET_KEY and ENCRYPTION_KEY (see below)&lt;/span&gt;
docker compose up &lt;span class="nt"&gt;-d&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the whole thing. Compose pulls the images, starts four containers, and runs the database migrations on first boot. Give it a minute, then open:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;http://localhost:3000&lt;/code&gt;&lt;/strong&gt; — the app (frontend)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;http://localhost:8000&lt;/code&gt;&lt;/strong&gt; — the API&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Create your account on the first-run screen and you're in. Your &lt;a href="https://datanika.io/docs/getting-started/" rel="noopener noreferrer"&gt;first pipeline&lt;/a&gt; — a source, a destination, a run — takes about five more minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's actually running
&lt;/h2&gt;

&lt;p&gt;Four containers, and it's worth knowing what each one does before you put it in production:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service&lt;/th&gt;
&lt;th&gt;Image&lt;/th&gt;
&lt;th&gt;Port&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;app&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;datanika&lt;/td&gt;
&lt;td&gt;3000, 8000&lt;/td&gt;
&lt;td&gt;The Reflex app — frontend + API in one process&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;celery&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;datanika&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Background worker: runs your extracts, loads, and dbt builds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;postgres&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;postgres:16&lt;/td&gt;
&lt;td&gt;5432&lt;/td&gt;
&lt;td&gt;Application database (your orgs, connections, run history — &lt;strong&gt;not&lt;/strong&gt; your warehouse)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;redis&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;redis:7&lt;/td&gt;
&lt;td&gt;6379&lt;/td&gt;
&lt;td&gt;Task broker for Celery + cache&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One thing that trips people up: &lt;strong&gt;&lt;code&gt;postgres&lt;/code&gt; here is Datanika's own metadata database, not your data warehouse.&lt;/strong&gt; Your extracted data lands wherever you point it — BigQuery, Snowflake, a separate Postgres, DuckDB on the same box. This container just holds Datanika's bookkeeping.&lt;/p&gt;

&lt;h2&gt;
  
  
  The configuration that actually matters
&lt;/h2&gt;

&lt;p&gt;Most of &lt;code&gt;.env&lt;/code&gt; has sane defaults. Two variables you &lt;strong&gt;must&lt;/strong&gt; set to real random values before anything touches production:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# JWT signing key — anyone who has this can forge login tokens&lt;/span&gt;
&lt;span class="nv"&gt;SECRET_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;openssl rand &lt;span class="nt"&gt;-hex&lt;/span&gt; 32&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="c"&gt;# Fernet key — encrypts every stored connector credential at rest&lt;/span&gt;
&lt;span class="nv"&gt;ENCRYPTION_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;ENCRYPTION_KEY&lt;/code&gt; is the one to guard: it's the Fernet key Datanika uses to encrypt every source and destination credential in the metadata DB. &lt;strong&gt;If you lose it, every stored credential becomes unrecoverable and you'll re-enter them all.&lt;/strong&gt; Back it up somewhere that isn't the same box.&lt;/p&gt;

&lt;p&gt;The rest, with their defaults:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variable&lt;/th&gt;
&lt;th&gt;What it's for&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;DATABASE_URL&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Metadata Postgres connection&lt;/td&gt;
&lt;td&gt;&lt;code&gt;postgresql+asyncpg://datanika:datanika@postgres:5432/datanika&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;REDIS_URL&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Celery broker&lt;/td&gt;
&lt;td&gt;&lt;code&gt;redis://redis:6379/0&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;SMTP_HOST&lt;/code&gt; / &lt;code&gt;SMTP_PORT&lt;/code&gt; / &lt;code&gt;SMTP_USER&lt;/code&gt; / &lt;code&gt;SMTP_PASSWORD&lt;/code&gt; / &lt;code&gt;EMAIL_FROM&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Email for run-failure alerts and invites&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;RECAPTCHA_SITE_KEY&lt;/code&gt; / &lt;code&gt;RECAPTCHA_SECRET_KEY&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Bot protection on signup (leave empty to disable)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Configure SMTP if you want failure notifications — a self-hosted pipeline that fails silently is worse than no pipeline. Everything else can wait.&lt;/p&gt;

&lt;h2&gt;
  
  
  Migrations and first boot
&lt;/h2&gt;

&lt;p&gt;Migrations run automatically the first time the &lt;code&gt;app&lt;/code&gt; container starts. If you ever need to run them by hand — say after pulling a new version — it's:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker compose &lt;span class="nb"&gt;exec &lt;/span&gt;app alembic upgrade &lt;span class="nb"&gt;head&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Taking it to production (the honest part)
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;docker compose up -d&lt;/code&gt; gets you a working instance. It does &lt;strong&gt;not&lt;/strong&gt; get you a production-grade one. Here's the checklist we'd actually run through, in order of how much it'll hurt to skip:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Put a reverse proxy in front and terminate TLS.&lt;/strong&gt; Nginx or &lt;a href="https://caddyserver.com/" rel="noopener noreferrer"&gt;Caddy&lt;/a&gt; (Caddy does automatic Let's Encrypt certs in about four lines). Bind the app containers to &lt;code&gt;127.0.0.1&lt;/code&gt; and let the proxy be the only thing on &lt;code&gt;:443&lt;/code&gt;. Never expose &lt;code&gt;:3000&lt;/code&gt;/&lt;code&gt;:8000&lt;/code&gt; to the internet directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Back up the metadata Postgres.&lt;/strong&gt; A nightly &lt;code&gt;pg_dump&lt;/code&gt; is the difference between "restore in ten minutes" and "rebuild every connection by hand." Ship it off the box:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   docker compose &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-T&lt;/span&gt; postgres pg_dump &lt;span class="nt"&gt;-U&lt;/span&gt; datanika datanika | &lt;span class="nb"&gt;gzip&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; datanika-&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%F&lt;span class="si"&gt;)&lt;/span&gt;.sql.gz
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Back up &lt;code&gt;ENCRYPTION_KEY&lt;/code&gt; alongside it — a database dump full of credentials you can no longer decrypt is not a backup.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Lock down the network.&lt;/strong&gt; Change the default Postgres/Redis passwords, restrict container ports to localhost, and put a firewall (ufw / security group) in front. Redis with no password on a public interface is a classic way to get owned.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Point monitoring at it.&lt;/strong&gt; The Compose file ships optional Grafana + Prometheus profiles; wire them up or point your existing stack at the app's metrics. You want to know a pipeline is failing before your stakeholders tell you the dashboard is stale — the &lt;a href="https://datanika.io/blog/slack-alerts-pipeline-failures/" rel="noopener noreferrer"&gt;Slack-alerts setup&lt;/a&gt; is the fastest win here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Right-size the box.&lt;/strong&gt; 8 GB RAM / 4 vCPU is a comfortable floor for real workloads. Extract jobs are memory-hungry in bursts; Celery is where that shows up.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of this is Datanika-specific — it's the standard "I now run a stateful service" checklist. But it's real work, and it's the honest cost of the $0 license.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bring your own database
&lt;/h2&gt;

&lt;p&gt;The bundled Postgres and Redis are convenient for getting started and fine for a small single-box deploy. For anything you care about, point Datanika at managed instances instead — set &lt;code&gt;DATABASE_URL&lt;/code&gt; and &lt;code&gt;REDIS_URL&lt;/code&gt; to your managed endpoints and the bundled containers become dead weight you can remove from the Compose file. Managed Postgres gets you backups, failover, and point-in-time recovery without you building any of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Upgrading
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;datanika-core
git pull origin master
docker compose up &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--build&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;app&lt;/code&gt; container runs &lt;code&gt;alembic upgrade head&lt;/code&gt; on startup, so migrations apply themselves. Pin to a tagged release rather than tracking &lt;code&gt;master&lt;/code&gt; if you want change control — and take a &lt;code&gt;pg_dump&lt;/code&gt; before every upgrade, because "roll back the database" is a lot easier than "figure out what the half-applied migration did."&lt;/p&gt;

&lt;h2&gt;
  
  
  Kubernetes, if that's your world
&lt;/h2&gt;

&lt;p&gt;A minimal Helm chart ships in-tree at &lt;code&gt;deploy/helm/datanika/&lt;/code&gt; — same image, one &lt;code&gt;app&lt;/code&gt; Deployment, one &lt;code&gt;celery&lt;/code&gt; Deployment, optional ingress. A few sharp edges to know before you &lt;code&gt;helm install&lt;/code&gt;: it needs a &lt;strong&gt;ReadWriteMany&lt;/strong&gt; storage class (the &lt;code&gt;app&lt;/code&gt; and &lt;code&gt;celery&lt;/code&gt; pods share a &lt;code&gt;dbt_projects&lt;/code&gt; volume), migrations run on every pod start (so keep &lt;code&gt;app.replicaCount=1&lt;/code&gt; until an HA migration hook lands), and you should disable the bundled single-replica Postgres/Redis in favor of managed ones. The &lt;a href="https://datanika.io/docs/self-hosting/" rel="noopener noreferrer"&gt;self-hosting docs&lt;/a&gt; have the full &lt;code&gt;values.yaml&lt;/code&gt; walkthrough.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-hosted vs. managed — the honest split
&lt;/h2&gt;

&lt;p&gt;Everything in the product is in the open-source core. What you're &lt;em&gt;not&lt;/em&gt; getting by self-hosting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Billing / metering&lt;/strong&gt; (Paddle integration) — irrelevant unless you're reselling&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Managed infrastructure and automatic updates&lt;/strong&gt; — you run &lt;code&gt;git pull&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Priority support with an SLA&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SSO (SAML/OIDC)&lt;/strong&gt; — gated to the Enterprise plan on managed cloud, though the SSO code itself lives in the open-source core&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If none of those matter to you, self-hosting isn't a downgrade — it's the same platform on your terms. This blog, and the pipelines behind it, run on exactly this stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to go next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/docs/self-hosting/" rel="noopener noreferrer"&gt;Self-hosting docs&lt;/a&gt;&lt;/strong&gt; — the canonical reference: full env-var table, Helm &lt;code&gt;values.yaml&lt;/code&gt;, production notes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/docs/getting-started/" rel="noopener noreferrer"&gt;Getting Started&lt;/a&gt;&lt;/strong&gt; — your first source → destination → run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/docs/architecture/" rel="noopener noreferrer"&gt;Architecture&lt;/a&gt;&lt;/strong&gt; — how &lt;code&gt;dlt&lt;/code&gt;, &lt;code&gt;dbt-core&lt;/code&gt;, Celery, and Reflex fit together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/connectors/" rel="noopener noreferrer"&gt;Browse all connectors&lt;/a&gt;&lt;/strong&gt; — 31 sources and 9 destinations, all included, no plan gating.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://datanika.io/blog/saas-12-euros/" rel="noopener noreferrer"&gt;The €12/mo stack, as it stood in April 2026&lt;/a&gt;&lt;/strong&gt; — the bill for running real software on one small VPS.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Clone it, run &lt;code&gt;docker compose up -d&lt;/code&gt;, and you own your data pipeline stack end to end. Or if you'd rather we run it — &lt;a href="https://app.datanika.io/" rel="noopener noreferrer"&gt;the managed free tier&lt;/a&gt; is one click and includes 10 GB/month.&lt;/p&gt;

</description>
      <category>tutorial</category>
      <category>selfhosting</category>
      <category>docker</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Four Identical Red Runs: One Was Our Watchdog Working, Three Were Its Corpse</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Tue, 15 Sep 2026 22:47:57 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/four-identical-red-runs-one-was-our-watchdog-working-three-were-its-corpse-3b60</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/four-identical-red-runs-one-was-our-watchdog-working-three-were-its-corpse-3b60</guid>
      <description>&lt;p&gt;A cron in one of our repositories stopped firing on 21 June and nobody noticed for ten weeks. Its only job was to rebuild this site each morning so that blog posts whose publish date had arrived would actually appear. So for ten weeks, scheduled posts did not publish. The failure mode of a cron is silence, and silence is the one thing no dashboard renders.&lt;/p&gt;

&lt;p&gt;We built a watchdog for it. The watchdog has four scheduled runs in its history and all four are red.&lt;/p&gt;

&lt;p&gt;Here is what makes this worth writing down: &lt;strong&gt;the first red was the watchdog working exactly as designed&lt;/strong&gt;, and the other three were it dying before it checked anything. From the outside they are indistinguishable — not merely the same colour, but the same run conclusion and the same step-by-step breakdown, line for line.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four runs
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;33330840430  2026-08-30T19:25:44Z  failure
33442245077  2026-08-31T21:36:52Z  failure
33550257729  2026-09-01T19:34:08Z  failure
33673601655  2026-09-02T19:29:21Z  failure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Expand any of them and you get the same two lines that matter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;step 5  Check every scheduled workflow in both public repos  = success
step 6  File an issue when a schedule has stopped            = failure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is run 1. It is also run 2, run 3 and run 4. We checked all four against the API rather than trusting the screen; they agree to the character.&lt;/p&gt;

&lt;p&gt;On 30 August, step 6 was red because the watchdog &lt;strong&gt;had found a stopped cron and was filing an issue about it&lt;/strong&gt;. The issue exists, machine-authored, timestamped 19:26:02Z, and it is correct in every particular:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;datanika-landing :: .github/workflows/daily-rebuild.yml&lt;/code&gt; last ran on a &lt;code&gt;schedule&lt;/code&gt; event at 2026-06-21T09:54:45+00:00 — 70.4 days ago. Its cron (&lt;code&gt;0 6 * * *&lt;/code&gt;) should have fired within 38h. It is &lt;code&gt;active&lt;/code&gt;, so this is NOT the 60-day disable; something else is stopping it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;On 31 August, 1 September and 2 September, step 6 was red because the watchdog had crashed in step 5 and there was nothing to report. It verified nothing on any of those nights.&lt;/p&gt;

&lt;h2&gt;
  
  
  What broke it, and it was not a commit
&lt;/h2&gt;

&lt;p&gt;Nothing in our repository changed between 30 and 31 August. What changed was a repository &lt;em&gt;setting&lt;/em&gt;: Dependabot became active on the repo. And when Dependabot is on, GitHub starts listing an extra workflow.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;gh api repos/OWNER/REPO/actions/workflows &lt;span class="nt"&gt;--jq&lt;/span&gt; &lt;span class="s1"&gt;'.workflows[].path'&lt;/span&gt;
.github/workflows/ci.yml
.github/workflows/deploy-pointer.yml
...
dynamic/dependabot/update-graph
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last one has &lt;code&gt;state: active&lt;/code&gt; and a display name of "Dependency Graph". It is not a file. We did not write it, it is not in the tree, and it is not in any branch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;gh&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;api&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"repos/OWNER/REPO/contents/dynamic/dependabot/update-graph?ref=master"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Not Found"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"documentation_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://docs.github.com/rest/repos/contents#get-repository-content"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"404"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;GitHub synthesises it. The workflows API returns it; the contents API has never heard of it. It is not documented as an exception anywhere we could find, and it appears in &lt;strong&gt;both&lt;/strong&gt; of our public repositories.&lt;/p&gt;

&lt;p&gt;Our watchdog walks every workflow the API returns and reads each one's YAML off the default branch to extract its cron:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_gh&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gh&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_parse_crons&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;repo&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;_gh&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;api&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;repos/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;repo&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/contents/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;?ref=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ref&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--jq&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;check=True&lt;/code&gt; turns the 404 into &lt;code&gt;CalledProcessError&lt;/code&gt;. Nothing catches it. The script dies mid-collection — and because it collects the repos in order, it died on the first one and never reached the second, which is &lt;em&gt;the repo the watchdog was built to watch&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;If you have any tooling that enumerates workflows through the API and then reads their files, run this against your repos now:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gh api repos/OWNER/REPO/actions/workflows &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--jq&lt;/span&gt; &lt;span class="s1"&gt;'.workflows[] | select(.path | startswith(".github/workflows/") | not) | .path'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Anything it prints will 404 on &lt;code&gt;contents/&lt;/code&gt;. &lt;code&gt;dynamic/pages/pages-build-deployment&lt;/code&gt; shows up the same way once GitHub Pages is enabled, so this is a family, not a one-off.&lt;/p&gt;

&lt;p&gt;The fix is a whitelist, not a &lt;code&gt;dynamic/&lt;/code&gt; blacklist — GitHub only ever executes workflows out of &lt;code&gt;.github/workflows/&lt;/code&gt;, so anything outside that path cannot be a workflow you own, and whatever prefix GitHub invents next year is handled without a code change.&lt;/p&gt;

&lt;p&gt;It specifically must &lt;strong&gt;not&lt;/strong&gt; become "swallow the 404". A 404 on a real &lt;code&gt;.github/workflows/*.yml&lt;/code&gt; means a file you were asked to check is unreadable — a token scope, a rename, an API change — and that has to stay fatal. Turning it into "no crons found" would make the watchdog report health from an error.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that cost us three nights
&lt;/h2&gt;

&lt;p&gt;The bug above is a fifteen-minute fix. The reason it survived three nights is a design problem, and it is the transferable half.&lt;/p&gt;

&lt;p&gt;Look at the reporting step. Every path through it ends the same way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; problems.md &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"::error::The watchdog exited non-zero but wrote no problems.md."&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"::error::That means it failed to RUN, not that it found a fault."&lt;/span&gt;
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi

if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$existing&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;gh issue comment &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$existing&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--body-file&lt;/span&gt; /tmp/comment.md
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi

&lt;/span&gt;gh issue create &lt;span class="nt"&gt;--title&lt;/span&gt; &lt;span class="s2"&gt;"Scheduled workflow stopped firing"&lt;/span&gt; &lt;span class="nt"&gt;--body-file&lt;/span&gt; /tmp/body.md
&lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three outcomes — &lt;em&gt;filed a new finding&lt;/em&gt;, &lt;em&gt;added to an existing finding&lt;/em&gt;, &lt;em&gt;could not run at all&lt;/em&gt; — collapsed into one signal. The run had to be red for the first two, because that is how a monitor gets your attention. So the third inherited the same colour, and the watchdog's own catastrophic failure was camouflaged by its success case.&lt;/p&gt;

&lt;p&gt;We have written before about &lt;a href="https://datanika.io/blog/github-actions-pipefail-exit-code/" rel="noopener noreferrer"&gt;a green that proves nothing&lt;/a&gt; and about &lt;a href="https://datanika.io/blog/alerts-that-could-never-fire/" rel="noopener noreferrer"&gt;alerts that could not fire at all&lt;/a&gt;, and about the inverse case, &lt;a href="https://datanika.io/blog/a-red-that-proves-nothing/" rel="noopener noreferrer"&gt;a red that means "I found nothing"&lt;/a&gt;. This is a fourth shape and it is the meanest of them, because the signal is &lt;em&gt;not&lt;/em&gt; useless — it is genuinely informative, one night in four. It just does not carry which thing it means.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A monitor has three states, not two: ran and clean, ran and found something, did not run.&lt;/strong&gt; If two of those share a colour, the pair that shares it is the pair you will conflate — and you will conflate it in the direction of "working", because that is the reading that requires no action.&lt;/p&gt;

&lt;h2&gt;
  
  
  The diagnostics were perfect and nobody read them
&lt;/h2&gt;

&lt;p&gt;This is the detail that stings. Open the log of any of the three dead runs and the workflow tells you, in plain English, exactly what happened:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;##[error]The watchdog exited non-zero but wrote no problems.md.
##[error]That means it failed to RUN, not that it found a fault.
##[error]Check the step above -- most likely the token cannot
##[error]read one of the repos, or the workflows API changed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Whoever wrote that anticipated this precise failure and left a message that names the right layer and distinguishes the two cases the conclusion collapses. It was there all three nights. It went unread, because a red tick on a default branch that carries other standing reds is not a thing anyone clicks into.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Diagnostics one level below the signal are not diagnostics.&lt;/strong&gt; They are an artifact you will find during the postmortem and feel bad about. If the message needs to be read, it has to travel through the same channel as the finding it disambiguates.&lt;/p&gt;

&lt;p&gt;So that is what the fix does. The watchdog now files an issue about &lt;em&gt;itself&lt;/em&gt; when it cannot run, under a deliberately different title, with a line stating what its own silence is now worth:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;While this issue is open, the absence of a "Scheduled workflow stopped firing" issue means nothing: the detector is down, not the crons proven up.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Both paths still exit non-zero — the workflow needs that — but the two conditions are now one glance apart in the issue list instead of one log dive apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  Twenty-six green tests, none of which touched the broken part
&lt;/h2&gt;

&lt;p&gt;The watchdog had a test suite. Twenty-six tests, green throughout all three dead nights.&lt;/p&gt;

&lt;p&gt;Every one of them exercised the comparator: given these crons and these last-run timestamps, which of them are overdue? That logic was never wrong. Not one test exercised &lt;em&gt;collection&lt;/em&gt; — the loop that asks the API what workflows exist and fetches each one's file. Collection was the thin shell around &lt;code&gt;gh api&lt;/code&gt; that the module's own docstring described as not the part worth testing.&lt;/p&gt;

&lt;p&gt;There is a second, quieter version of the same mistake in there. The suite did check that collection found the expected number of workflows — but it asserted a &lt;strong&gt;total&lt;/strong&gt; across both repos. Our first repo has enough workflows on its own to satisfy that total, so the count passed while the second repo was never reached. The fix asserts per repo, which is the assertion that would have failed.&lt;/p&gt;

&lt;p&gt;Whenever you decide some part of a program is too thin to test, you have made a claim about where failure lives. Write it down as a claim, because it is one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The last twist
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Update, 15 September 2026.&lt;/strong&gt; This section describes the morning of 3 September, when it was written, and it stopped being true that evening. The fix reached the default branch that day: the scheduled run that night completed cleanly, and so did the five after it. Since 9 September the watchdog has been red every night again, and not with the crash described here — it runs, and its most recent report names its own failed runs as the schedule that stopped. The original text follows as it was published.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;code&gt;schedule:&lt;/code&gt; triggers only ever run the copy of a workflow on the repository's &lt;strong&gt;default branch&lt;/strong&gt;. Our fix is merged to the integration branch and has not been promoted. Tonight's scheduled run will crash again, in exactly the way described above, and there is no branch we could put the fix on that would change that.&lt;/p&gt;

&lt;p&gt;Which means the fix cannot be proven by the mechanism it fixes. The only successful run in this watchdog's entire history is a &lt;code&gt;workflow_dispatch&lt;/code&gt; we triggered by hand against the fix branch. That proves the code runs. It does not prove the schedule fires, and it does not prove the schedule fires &lt;em&gt;this&lt;/em&gt; code.&lt;/p&gt;

&lt;p&gt;Our own workflow header had already warned about this trap, from the last time it bit us:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do not read a green run here as proof that it works. The thing under test is whether an &lt;em&gt;unattended&lt;/em&gt; &lt;code&gt;event=schedule&lt;/code&gt; run appears at all; a &lt;code&gt;workflow_dispatch&lt;/code&gt; proves only that the dispatch step is sound.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The query that answers it honestly filters on the event:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gh api &lt;span class="s2"&gt;"repos/OWNER/REPO/actions/workflows/NAME.yml/runs?per_page=100"&lt;/span&gt; &lt;span class="nt"&gt;--paginate&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--jq&lt;/span&gt; &lt;span class="s1"&gt;'.workflow_runs[] | select(.event=="schedule") | .created_at'&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run that against the cron this whole story started with, and the gap is unmissable: daily &lt;code&gt;schedule&lt;/code&gt; runs from 16 April to 21 June, then nothing until 31 August. Seventy-one days. During that window a hand-triggered dispatch went green — which is precisely why a green dispatch was never evidence.&lt;/p&gt;

&lt;p&gt;One honest footnote while we are counting. The remedy we applied for that cron was to move it off the top of the hour, on the theory that GitHub's queue is shortest away from &lt;code&gt;:00&lt;/code&gt;. Three scheduled runs since: 4h49m, 5h17m and 7h48m late. Three samples is not a refutation, but it is not the improvement we told ourselves we were buying either, and we would rather print that than quietly drop it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four things to check in your own repositories
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;List every workflow whose path is not under &lt;code&gt;.github/workflows/&lt;/code&gt;.&lt;/strong&gt; If anything prints, every tool you have that reads workflow files by path is one API call from a crash.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For each monitor you run, ask what a red means.&lt;/strong&gt; If "found a problem" and "could not run" share a conclusion, you own one signal doing two jobs, and it will fail in the direction that looks fine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Give the clean path something to emit.&lt;/strong&gt; A monitor that is only ever heard from when it has bad news is indistinguishable from one that has died — the useful assertion is on a heartbeat going stale, not on an alarm being absent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask when each scheduled workflow last ran on an &lt;code&gt;event=schedule&lt;/code&gt;&lt;/strong&gt;, not when it last ran. Manual and dispatched runs will happily paper over a cron that has been dead for two months.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then take the monitor you trust most and break it on purpose — not the system it watches, the monitor itself — and see whether anything looks different.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Datanika is an open-source data platform that runs dlt extract-and-load and dbt-core transformations behind one UI, with scheduling, run history and a REST API. Every workflow quoted here lives in our public repositories, defects included. It is AGPL-3.0 and self-hostable — see &lt;a href="https://datanika.io/docs/architecture/" rel="noopener noreferrer"&gt;the architecture&lt;/a&gt;, or &lt;a href="https://datanika.io/connectors/" rel="noopener noreferrer"&gt;browse the connectors&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ci</category>
      <category>githubactions</category>
      <category>monitoring</category>
      <category>observability</category>
    </item>
    <item>
      <title>The Check That Went Red for Doing Its Job</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Tue, 15 Sep 2026 13:26:07 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/the-check-that-went-red-for-doing-its-job-ek3</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/the-check-that-went-red-for-doing-its-job-ek3</guid>
      <description>&lt;p&gt;We have written twice recently about checks that could not fail — &lt;a href="https://datanika.io/blog/alerts-that-could-never-fire/" rel="noopener noreferrer"&gt;four alerts that could never fire&lt;/a&gt; and &lt;a href="https://datanika.io/blog/github-actions-pipefail-exit-code/" rel="noopener noreferrer"&gt;a nightly suite that passed for eight nights while twelve tests failed&lt;/a&gt;. Both are the same defect: a green signal that would have looked identical had the thing it watches been broken.&lt;/p&gt;

&lt;p&gt;This one is the mirror image, and it is worse in a way that took us a while to articulate.&lt;/p&gt;

&lt;h2&gt;
  
  
  One run, three levels, three answers
&lt;/h2&gt;

&lt;p&gt;Here is a single GitHub Actions run from one of our repositories, read three ways.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The run conclusion:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;failure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The two jobs in it:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="err"&gt;JOB&lt;/span&gt; &lt;span class="py"&gt;parity&lt;/span&gt;        &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;success&lt;/span&gt;
&lt;span class="err"&gt;JOB&lt;/span&gt; &lt;span class="py"&gt;config-fields&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;failure&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The steps inside the failing job:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 Set up job                                    = success
2 Run actions/checkout@v6                       = success
3 Fetch core's shipped connection schema        = success
4 Compare documented fields against the form    = success
5 File an issue on drift                        = failure      &amp;lt;-- the only red
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every step that does real work succeeded. The one that failed was the step that &lt;em&gt;reports&lt;/em&gt;. And here is its entire log:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Already tracked in #449; not filing again.
##[error]Process completed with exit code 1.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The check went red because it found nothing new to say.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug is four lines and looks completely reasonable
&lt;/h2&gt;

&lt;p&gt;The workflow compares two catalogues that live in different repositories and files an issue when they disagree. Filing the same issue every morning would be useless, so it suppresses duplicates:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;existing&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;gh issue list &lt;span class="nt"&gt;--state&lt;/span&gt; open &lt;span class="nt"&gt;--search&lt;/span&gt; &lt;span class="s2"&gt;"in:title Connector config field drift"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--json&lt;/span&gt; number &lt;span class="nt"&gt;--jq&lt;/span&gt; &lt;span class="s1"&gt;'.[0].number // empty'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$existing&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Already tracked in #&lt;/span&gt;&lt;span class="nv"&gt;$existing&lt;/span&gt;&lt;span class="s2"&gt;; not filing again."&lt;/span&gt;
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read that in review and nothing jumps out. &lt;code&gt;exit 1&lt;/code&gt; after "we found drift" feels right — drift &lt;em&gt;is&lt;/em&gt; a problem, and a red check is how a cron gets anyone's attention.&lt;/p&gt;

&lt;p&gt;But this branch is not "we found drift." It is &lt;strong&gt;"we found drift, and it is the same drift as yesterday, already filed, already assigned."&lt;/strong&gt; Nothing changed. Nothing needs doing. And because the issue will be open for as long as it takes to fix a 36-page documentation mismatch, this branch runs &lt;strong&gt;every day until then&lt;/strong&gt;, painting the check red every time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a permanent red is worse than a permanent green
&lt;/h2&gt;

&lt;p&gt;This is the part worth taking away, and it is not symmetric with the "green that proves nothing" story.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A green that proves nothing &lt;strong&gt;fails to inform&lt;/strong&gt;. A red that repeats forever &lt;strong&gt;destroys the channel.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A useless green leaves you no worse off than having no check. A permanent red actively trains every human near the repository to stop reading that signal — and the steps above it in the same job are the reason the cron exists. In our case those steps fetch a schema from another repository and parse a data file by regex. Both are exactly the kind of thing that breaks quietly when someone renames a field. The job even has explicit guards for it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;::error::Parsed $slugs slugs out of connectors.ts — the file shape changed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That guard reports through the channel the duplicate-suppression branch was jamming. Had the file shape actually changed, the message would have arrived on a check that everyone had already learned to ignore, and it would have looked exactly like yesterday.&lt;/p&gt;

&lt;p&gt;There is a second cost, which we hit within the hour. The red job sat next to a &lt;em&gt;succeeding&lt;/em&gt; one and a &lt;em&gt;stale&lt;/em&gt; auto-filed issue titled "Connector count drift." Glanced at, the repository said the connector count was broken. It was not — the counts agree, on both sides, in the workflow's own log, and on the live site. We nearly spent a session re-fixing something that was already fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The line we drew
&lt;/h2&gt;

&lt;p&gt;The distinction that resolves it is not "is there a problem" but &lt;strong&gt;"did anything change that needs a human?"&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;branch&lt;/th&gt;
&lt;th&gt;exit&lt;/th&gt;
&lt;th&gt;reasoning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;filed a &lt;strong&gt;new&lt;/strong&gt; issue&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;something changed today; a red run is a standing signal that is easy to see&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;an issue &lt;strong&gt;already exists&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;nothing changed, nothing to do, and the open issue is already the tracking mechanism&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Note that we did not simply make the cron always green. The &lt;code&gt;exit 1&lt;/code&gt; after a genuine &lt;code&gt;gh issue create&lt;/code&gt; stays exactly as it was. A rule that removes the check's ability to ever go red would be the original defect arriving from the other direction, so the test we added asserts the narrow version &lt;em&gt;and&lt;/em&gt; asserts that both jobs still exit non-zero after really filing something.&lt;/p&gt;

&lt;h2&gt;
  
  
  Read the level that answers your question
&lt;/h2&gt;

&lt;p&gt;The same run demonstrates a second thing, and it is free.&lt;/p&gt;

&lt;p&gt;Three conclusions were available — run, job, step — and they were &lt;code&gt;failure&lt;/code&gt;, &lt;code&gt;success&lt;/code&gt;/&lt;code&gt;failure&lt;/code&gt;, and one red step among five. &lt;strong&gt;Every one is a true statement about a different question.&lt;/strong&gt; If you ask the run whether your drift check is healthy, the answer is a confident no, for a reason that has nothing to do with drift.&lt;/p&gt;

&lt;p&gt;Fetch step-level data whenever the question is about a particular thing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gh api repos/&amp;lt;owner&amp;gt;/&amp;lt;repo&amp;gt;/actions/runs/&amp;lt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/jobs &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--jq&lt;/span&gt; &lt;span class="s1"&gt;'.jobs[] | {name} + {steps: [.steps[] | {number, name, conclusion}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two related traps we have hit: a job's own conclusion can be red because of an artifact upload rather than anything it tested, and a &lt;em&gt;fetched log&lt;/em&gt; can silently omit steps entirely while the API reports them as &lt;code&gt;success&lt;/code&gt;. &lt;strong&gt;Read outcomes from the API, and never conclude from an absence.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Three checks you can run on your own repositories
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Grep your automation for &lt;code&gt;exit 1&lt;/code&gt; in a de-duplication or "already handled" branch.&lt;/strong&gt; Anywhere a script's happy path for "nothing new" is an error exit. This includes retry guards, lockfile checks, and "skip because it is already deployed".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;List the checks that have been red for more than a week.&lt;/strong&gt; For each, ask what would happen if a &lt;em&gt;different&lt;/em&gt; thing in that job broke tomorrow. If the answer is "the colour would not change", the job has no remaining signal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For every scheduled workflow, ask what its steady state should be.&lt;/strong&gt; A cron that watches for drift is supposed to be green almost always and red on the day something moves. If yours is red most days by design, the design is wrong.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  One note on writing the test
&lt;/h2&gt;

&lt;p&gt;We pinned the fix with a test that parses the workflow file, and the first run of its mutation harness reported:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;M1 :: MUTATION DID NOT APPLY
M2 :: MUTATION DID NOT APPLY
M3 :: MUTATION DID NOT APPLY
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The working copy of the workflow is CRLF and the harness's multi-line anchors used bare &lt;code&gt;\n&lt;/code&gt;, so they matched nothing. That is worth stating because of what the alternative looks like: a mutation tool that silently changes nothing will report every mutation as "the test caught it", and you will believe you have a proven guard when you have an untested one. &lt;strong&gt;Make a harness assert that its mutation actually landed&lt;/strong&gt; — we compare the file's hash before and after — and treat "did not apply" as a distinct outcome from "passed".&lt;/p&gt;

&lt;p&gt;With the anchors corrected, all four mutations went red: both duplicate-suppression branches restored to &lt;code&gt;exit 1&lt;/code&gt;, the scope control that deletes a legitimate &lt;code&gt;exit 1&lt;/code&gt;, and a positive control that renames the log line the matcher keys on.&lt;/p&gt;




&lt;p&gt;Datanika is an open-source data platform — extraction, loading, transformation and scheduling in one UI. The workflows described here are in our public repositories, defects included. &lt;a href="https://app.datanika.io" rel="noopener noreferrer"&gt;Try it free&lt;/a&gt;, or &lt;a href="https://datanika.io/docs/self-hosting/" rel="noopener noreferrer"&gt;read the self-hosting guide&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ci</category>
      <category>githubactions</category>
      <category>monitoring</category>
      <category>testing</category>
    </item>
    <item>
      <title>Four Alerts That Could Never Fire, and How We Found Them</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Tue, 15 Sep 2026 13:21:05 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/four-alerts-that-could-never-fire-and-how-we-found-them-5gh8</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/four-alerts-that-could-never-fire-and-how-we-found-them-5gh8</guid>
      <description>&lt;p&gt;We spent a week auditing our own alerting, and the useful question turned out not to be &lt;em&gt;"is anything alerting?"&lt;/em&gt; It was &lt;strong&gt;"would this alert look any different if the thing it watches had failed?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For four of them the answer was no. Not "it fires late", not "the threshold is loose" — these could not produce a red signal at all, in the exact circumstance each was written for. Every dashboard was green the entire time, and the green was not evidence of anything.&lt;/p&gt;

&lt;p&gt;Here they are with the mechanisms, because each one is a shape you can go and check for in your own stack in about ten minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The rule that could only fire while the system was healthy
&lt;/h2&gt;

&lt;p&gt;We wanted an alert for &lt;em&gt;"a scheduled maintenance task has stopped arriving."&lt;/em&gt; The natural expression:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;increase(celery_tasks_total{task="datanika.run_maintenance"}[2h]) &amp;lt; 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read it in English and it is obviously right: fire when fewer than one run happened in two hours. It evaluates cleanly. It shows up in the rule list. Its health reads &lt;code&gt;ok&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It cannot fire.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In PromQL, a bare comparison between a vector and a scalar is a &lt;em&gt;filter&lt;/em&gt;, not a boolean.&lt;/strong&gt; &lt;code&gt;X &amp;lt; 1&lt;/code&gt; does not return true or false. It returns &lt;em&gt;&lt;code&gt;X&lt;/code&gt;'s own value&lt;/em&gt;, for every series where the condition holds, and drops the rest. So the number that reaches the alerting threshold is not &lt;code&gt;1&lt;/code&gt; and not &lt;code&gt;true&lt;/code&gt; — it is whatever &lt;code&gt;increase()&lt;/code&gt; computed, which by the definition of the filter is &lt;strong&gt;strictly less than 1&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Our rules pair that with a Grafana threshold of &lt;code&gt;gt [0]&lt;/code&gt;: fire when the reduced value is greater than zero. Now do the arithmetic for the case the rule exists to catch. A task that has completely stopped arriving has &lt;code&gt;increase(...) == 0&lt;/code&gt;. Zero passes the &lt;code&gt;&amp;lt; 1&lt;/code&gt; filter, so a series &lt;em&gt;is&lt;/em&gt; produced — and then &lt;code&gt;0 &amp;gt; 0&lt;/code&gt; is &lt;strong&gt;false&lt;/strong&gt;, so nothing fires.&lt;/p&gt;

&lt;p&gt;The rule is live only in the open interval &lt;code&gt;(0, 1)&lt;/code&gt;. It detects a task that has partially stopped and is structurally blind to one that has entirely stopped.&lt;/p&gt;

&lt;p&gt;Demonstrated against the engine rather than argued:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;vector(0) &amp;lt; 1        -&amp;gt; series=1 value=0     # gt [0] on 0 is FALSE -&amp;gt; dead rule
vector(0) &amp;lt; bool 1   -&amp;gt; series=1 value=1     # gt [0] on 1 is TRUE  -&amp;gt; fires
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;bool&lt;/code&gt; is the fix: it converts the filter into the 1/0 comparison everyone assumed they were writing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The part worth stealing is how it was found.&lt;/strong&gt; The engineer writing that rule used &lt;code&gt;&amp;lt; bool 1&lt;/code&gt; deliberately, and then &lt;em&gt;mutated their own correct code&lt;/em&gt; to the naive &lt;code&gt;&amp;lt; 1&lt;/code&gt; to check whether the linter would have caught the mistake had they made it.&lt;/p&gt;

&lt;p&gt;It did not. &lt;strong&gt;196 tests passed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Four other mutations run the same way — a dropped scrape job, a dropped deploy step, a drifted subquery step, a bare staleness comparison — all went correctly red, which is what proves the harness was armed and this one check simply does not look. A satisfiability check existed; it recognised the &lt;code&gt;== N&lt;/code&gt; shape and skipped silently on everything else, and a skip is a pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The exporter that was scraped, and received nothing
&lt;/h2&gt;

&lt;p&gt;Our task metrics were collected by code that had been in the repo since the start:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;celery_tasks_total&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;labels&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;failure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There was an alert on that counter. It had never fired. There had also never been a task failure it should have caught, so nobody thought about it.&lt;/p&gt;

&lt;p&gt;Two things were wrong, and each on its own is fatal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The counter incremented in one process and was served from another.&lt;/strong&gt; &lt;code&gt;celery_tasks_total&lt;/code&gt; is incremented inside the &lt;strong&gt;Celery worker&lt;/strong&gt;. &lt;code&gt;/metrics&lt;/code&gt; is a route in the &lt;strong&gt;web&lt;/strong&gt; process. Two processes, two &lt;code&gt;prometheus_client&lt;/code&gt; registries, no shared state. The counter the worker maintains is not in the payload the web process serves, and never was.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And Prometheus was not scraping Celery at all.&lt;/strong&gt; Seven targets configured, none of them the worker. So even a correctly located counter had no path to the time series database.&lt;/p&gt;

&lt;p&gt;The alert's query returned &lt;strong&gt;zero series&lt;/strong&gt;, forever. And here is the part that makes this a false-negative rather than a visible outage: our alert rules run with &lt;code&gt;noDataState: OK&lt;/code&gt;, which is deliberate and correct for filtering expressions — a healthy system genuinely produces no rows. So &lt;em&gt;"the metric does not exist"&lt;/em&gt; and &lt;em&gt;"nothing is wrong"&lt;/em&gt; arrive at the alert engine as the same signal.&lt;/p&gt;

&lt;p&gt;The general form, which we now have written down:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An exporter that is scraped but has stopped receiving events looks exactly like a quiet system.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The fix was not to move the counter. It was to stop asking the application to report on itself and read the &lt;strong&gt;broker's own event stream&lt;/strong&gt; instead, with a dedicated exporter and the worker started with &lt;code&gt;-E&lt;/code&gt;. Which introduced its own trap, so it goes in the list too: an exporter pointed at a worker &lt;em&gt;without&lt;/em&gt; &lt;code&gt;-E&lt;/code&gt; still emits worker-liveness metrics and zero task metrics — indistinguishable from a worker that simply has not run a task yet. You tell them apart by waiting for a scheduled firing and re-querying, never by reading the exporter's own health.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. &lt;code&gt;curl -sf&lt;/code&gt; succeeds on an HTML error page
&lt;/h2&gt;

&lt;p&gt;This one is the cheapest to reproduce and probably the most widespread.&lt;/p&gt;

&lt;p&gt;A verification runbook, gating a pricing change, contained this step:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sf&lt;/span&gt; https://app.example.com/metrics | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s2"&gt;"bytes_processed|bytes_quota"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under it, three checkboxes to tick when the metrics appear. Measured:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;http_code=200   size=5106   content_type=text/html;charset=utf-8
&lt;span class="cp"&gt;&amp;lt;!DOCTYPE html&amp;gt;&lt;/span&gt;&lt;span class="nt"&gt;&amp;lt;html&lt;/span&gt; &lt;span class="na"&gt;lang=&lt;/span&gt;&lt;span class="s"&gt;"en"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;…
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;/metrics&lt;/code&gt; had no entry in the reverse-proxy config. Our proxy routes a specific list of paths to the backend and sends &lt;strong&gt;everything else&lt;/strong&gt; to the single-page app, so an unrouted backend path does not 404 — it silently resolves to the frontend and returns the app shell with a cheerful &lt;strong&gt;200&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;curl -f&lt;/code&gt; fails on HTTP status codes at or above 400. A 200 carrying an HTML page is not a status error, so &lt;strong&gt;&lt;code&gt;-f&lt;/code&gt; does not trigger and &lt;code&gt;curl&lt;/code&gt; exits 0&lt;/strong&gt;. The pipe then hands &lt;code&gt;grep&lt;/code&gt; five kilobytes of HTML, &lt;code&gt;grep&lt;/code&gt; matches nothing, and the operator sees no output and no error.&lt;/p&gt;

&lt;p&gt;The checkboxes under that command could never be ticked, no matter what the application did. And the runbook's own troubleshooting table sent the reader to the wrong layer — &lt;em&gt;"no metrics after 2h → check the worker logs"&lt;/em&gt; — for a problem that was one line of proxy config.&lt;/p&gt;

&lt;p&gt;Two rules came out of it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;-f&lt;/code&gt; is a status check, not a content check.&lt;/strong&gt; If you are grepping a response, assert on the &lt;code&gt;Content-Type&lt;/code&gt; or on a string you know must be present, and fail the step when it is absent. &lt;code&gt;curl -sf … | grep -q 'expected'&lt;/code&gt; at least fails on the pipe's own exit status.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Any backend route outside your proxied prefixes needs its own entry, or it becomes your SPA.&lt;/strong&gt; The failure is silent, returns 200, and is invisible to anything that only checks status codes.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. The pager that rang for something that had not happened
&lt;/h2&gt;

&lt;p&gt;The other direction belongs in the same audit, because an instrument that manufactures incidents costs you the same credibility as one that hides them — and it burns it faster.&lt;/p&gt;

&lt;p&gt;Our end-to-end suite has an auto-filer: on failure it opens a tracker issue and pages. It fires on a &lt;strong&gt;job-level&lt;/strong&gt; &lt;code&gt;failure()&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;One night it filed and paged for a failed end-to-end run. Reading the step outcomes from the API rather than the log:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;conclusion&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Run gating E2E specs&lt;/td&gt;
&lt;td&gt;✅ success&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Detect flaky gating specs&lt;/td&gt;
&lt;td&gt;✅ success&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Run informational E2E specs&lt;/td&gt;
&lt;td&gt;✅ success&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assert the specs were actually collected&lt;/td&gt;
&lt;td&gt;✅ success&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Upload test report&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;🔴 &lt;strong&gt;failure&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Telegram alert on failure&lt;/td&gt;
&lt;td&gt;fired&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;File issue on failure&lt;/td&gt;
&lt;td&gt;fired&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The only error anywhere in the job:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Attempt 1..4 of 5 failed with error: Request timeout:
  /twirp/github.actions.results.api.v1.ArtifactService/CreateArtifact
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Every spec passed.&lt;/strong&gt; The job was red because an artifact upload timed out.&lt;/p&gt;

&lt;p&gt;What makes this worse than a cosmetic false alarm is what the filer writes into the issue it creates: a sentence stating that only the gating tier can open this report. That sentence became affirmatively false at the moment it was filed — and it is exactly the line a reader uses to decide how much to trust what they are reading. The tracker asserted a test failure that had not occurred, in a thread whose entire purpose is to be believed.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;if: failure()&lt;/code&gt; at job level means &lt;em&gt;"any step failed"&lt;/em&gt;. If you want &lt;em&gt;"the tests failed"&lt;/em&gt;, key the alert on the test step's own outcome:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run specs&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;specs&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pytest ...&lt;/span&gt;

&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Page on real failure&lt;/span&gt;
  &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;steps.specs.outcome == 'failure'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The pattern under all four
&lt;/h2&gt;

&lt;p&gt;Every one of these measured the &lt;strong&gt;instrument&lt;/strong&gt; rather than the system:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;what was green&lt;/th&gt;
&lt;th&gt;what it actually recorded&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;the alert rule's health&lt;/td&gt;
&lt;td&gt;that the expression parses and evaluates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the scrape target being &lt;code&gt;up&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;that an HTTP endpoint answered, not that it carried our data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;curl -sf&lt;/code&gt; exiting 0&lt;/td&gt;
&lt;td&gt;that some server returned a status below 400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the E2E job's red&lt;/td&gt;
&lt;td&gt;that some step in the job failed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None of them is broken in the sense of throwing an error. Each answers a real question accurately. The question is just not the one anybody thought was being asked.&lt;/p&gt;

&lt;p&gt;So the check we now run on every new alarm is a single sentence:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Would this signal look different if the thing it watches had failed?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And the only honest way to answer it is to make the thing fail. Not in a synthetic fixture built to satisfy the assertion — against the real artifact, in the real failure mode. Break the rule and watch it go red. Stop the worker and watch the metric vanish, then check what your &lt;code&gt;noDataState&lt;/code&gt; does with a vanished series. Point the &lt;code&gt;curl&lt;/code&gt; at a path you know is unrouted and confirm the step fails.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A passing check is not evidence until you have seen it fail.&lt;/strong&gt; We keep relearning that in a new costume roughly once a fortnight — most recently as &lt;a href="https://datanika.io/blog/github-actions-pipefail-exit-code/" rel="noopener noreferrer"&gt;a nightly CI job that reported success over twelve failing tests for eight nights&lt;/a&gt;, because a pipe to &lt;code&gt;tee&lt;/code&gt; discarded the exit code and GitHub Actions' default shell does not set &lt;code&gt;pipefail&lt;/code&gt;. Same family, different layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four things to go and check in your own stack
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Grep your alert expressions for a bare &lt;code&gt;&amp;lt;&lt;/code&gt; or &lt;code&gt;&amp;gt;&lt;/code&gt; against a scalar&lt;/strong&gt;, and check what value actually reaches your threshold evaluator. In PromQL, add &lt;code&gt;bool&lt;/code&gt; unless you specifically want the filter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For every counter you alert on, confirm the process that increments it is the process that serves it.&lt;/strong&gt; With any pre-fork or multi-process server, that is not automatic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Find every &lt;code&gt;curl -sf&lt;/code&gt; in your runbooks and CI&lt;/strong&gt;, and make each one assert on content, not just status. Then point one at a deliberately wrong path and confirm it fails.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read your alert conditions for scope.&lt;/strong&gt; &lt;code&gt;if: failure()&lt;/code&gt; and &lt;code&gt;on: failure&lt;/code&gt; are usually broader than the thing you meant.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then take one working alert and break the underlying system on purpose. If nothing turns red, you have found the fifth one.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Datanika is an open-source data platform that runs dlt extract-and-load and dbt-core transformations behind one UI, with scheduling, run history and a REST API. It is AGPL-3.0 and self-hostable with a single &lt;code&gt;docker compose up&lt;/code&gt; — see &lt;a href="https://datanika.io/docs/architecture/" rel="noopener noreferrer"&gt;the architecture&lt;/a&gt;, or &lt;a href="https://datanika.io/connectors/" rel="noopener noreferrer"&gt;browse the connectors&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>monitoring</category>
      <category>prometheus</category>
      <category>alerting</category>
      <category>observability</category>
    </item>
    <item>
      <title>Our Nightly CI Passed for Eight Nights While Twelve Tests Failed</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Wed, 09 Sep 2026 13:40:46 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/our-nightly-ci-passed-for-eight-nights-while-twelve-tests-failed-35l7</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/our-nightly-ci-passed-for-eight-nights-while-twelve-tests-failed-35l7</guid>
      <description>&lt;p&gt;Our nightly connector smoke suite reported &lt;code&gt;success&lt;/code&gt; eight nights running. Here is what it actually printed on each of those nights:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;12 failed, 9 passed, 4 warnings in 44.01s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Twelve of twenty-one probes failing, every night, for at least eight consecutive nights — the log retention window is the only reason the count stops at eight. The step reported &lt;code&gt;completed/success&lt;/code&gt;. The job reported &lt;code&gt;success&lt;/code&gt;. The workflow reported &lt;code&gt;success&lt;/code&gt;. The &lt;code&gt;Telegram alert on failure&lt;/code&gt; step, guarded by &lt;code&gt;if: failure()&lt;/code&gt;, was &lt;code&gt;skipped&lt;/code&gt; every single night, because &lt;code&gt;failure()&lt;/code&gt; never became true.&lt;/p&gt;

&lt;p&gt;Nobody had broken anything that week. The suite had been lying since long before.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mechanism is four words of shell
&lt;/h2&gt;

&lt;p&gt;Here is the step, near enough verbatim:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run connector smoke tests&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;pytest tests/test_connector_smoke/ -v --tb=short -rs | tee /tmp/smoke.log&lt;/span&gt;
    &lt;span class="s"&gt;if grep -qE '[0-9]+ skipped' /tmp/smoke.log; then&lt;/span&gt;
      &lt;span class="s"&gt;echo "::error::Smoke probes were SKIPPED"&lt;/span&gt;
      &lt;span class="s"&gt;exit 1&lt;/span&gt;
    &lt;span class="s"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A shell pipeline's exit status is the exit status of its &lt;strong&gt;last&lt;/strong&gt; command. The last command here is &lt;code&gt;tee&lt;/code&gt;, and &lt;code&gt;tee&lt;/code&gt; succeeded — it wrote the file it was asked to write. &lt;code&gt;pytest&lt;/code&gt; exited non-zero into a void.&lt;/p&gt;

&lt;p&gt;That much is ordinary POSIX shell, and most people who have written a CI pipeline know it. The part that catches you is the second half.&lt;/p&gt;

&lt;h2&gt;
  
  
  GitHub Actions gives you &lt;code&gt;pipefail&lt;/code&gt; only if you ask for it, and asking looks like a no-op
&lt;/h2&gt;

&lt;p&gt;Bash has a flag for exactly this. &lt;code&gt;set -o pipefail&lt;/code&gt; makes a pipeline return the rightmost non-zero exit status, so &lt;code&gt;pytest ... | tee&lt;/code&gt; fails when &lt;code&gt;pytest&lt;/code&gt; fails.&lt;/p&gt;

&lt;p&gt;GitHub Actions runs &lt;code&gt;run:&lt;/code&gt; steps on Linux and macOS with a default shell of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;bash &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;0&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;-e&lt;/code&gt; but no &lt;code&gt;-o pipefail&lt;/code&gt;. Now write the step as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run connector smoke tests&lt;/span&gt;
  &lt;span class="na"&gt;shell&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;bash&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and the runner invokes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;bash &lt;span class="nt"&gt;--noprofile&lt;/span&gt; &lt;span class="nt"&gt;--norc&lt;/span&gt; &lt;span class="nt"&gt;-eo&lt;/span&gt; pipefail &lt;span class="o"&gt;{&lt;/span&gt;0&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Naming the shell you were already using turns on &lt;code&gt;pipefail&lt;/code&gt;. That is a real asymmetry in the product, it is documented, and it reads as a no-op in a diff. A reviewer looking at &lt;code&gt;shell: bash&lt;/code&gt; on a step that was already running bash sees tidying, not a behavioural change — which cuts both ways: it is easy to add without argument, and easy for someone to delete later as noise.&lt;/p&gt;

&lt;p&gt;We had &lt;strong&gt;no &lt;code&gt;shell:&lt;/code&gt;, no &lt;code&gt;defaults:&lt;/code&gt; and no &lt;code&gt;pipefail&lt;/code&gt;&lt;/strong&gt; anywhere in that workflow file. So the only thing left that could fail the step was the &lt;code&gt;grep&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part actually worth writing down
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;grep&lt;/code&gt; is not a mistake. It exists for a good reason and it was working.&lt;/p&gt;

&lt;p&gt;Our connector probes need live credentials. A probe with no credentials, or one whose client library will not import, used to &lt;em&gt;skip&lt;/em&gt; — and a suite that skips everything passes loudly while testing nothing. So we changed &lt;code&gt;conftest.py&lt;/code&gt; to convert both cases from &lt;strong&gt;skip&lt;/strong&gt; into &lt;strong&gt;fail&lt;/strong&gt;, on the reasoning that a healthy run should have zero skips, which makes "any skip at all" a usable alarm. Then we added the &lt;code&gt;grep&lt;/code&gt; to enforce it.&lt;/p&gt;

&lt;p&gt;Read those two decisions in isolation and both are right. Read them together and this falls out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the outcome the guard watches for — &lt;strong&gt;skip&lt;/strong&gt; — is now the rare one, by construction;&lt;/li&gt;
&lt;li&gt;the outcome that actually happens — &lt;strong&gt;fail&lt;/strong&gt; — is the one the pipe throws away.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The guard was aimed at the hole that existed &lt;em&gt;before&lt;/em&gt; the &lt;code&gt;conftest&lt;/code&gt; change. The &lt;code&gt;conftest&lt;/code&gt; change moved every real failure into the blind spot. No commit introduced the bug; the second correct change walked the failure mode into the first correct change's shadow.&lt;/p&gt;

&lt;p&gt;That is the shape to look for, and it is not rare. When you tighten one behaviour, the checks written against the old behaviour do not fail — they go quiet, and quiet and healthy look identical from outside.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was underneath
&lt;/h2&gt;

&lt;p&gt;The twelve failures were all real, and none was caused by this bug. A couple of trial accounts had lapsed and were returning &lt;code&gt;401&lt;/code&gt; and &lt;code&gt;403&lt;/code&gt;. A message-broker probe was pointed at a cluster that no longer matched the one our credentials were minted against. Several probes were missing environment variables entirely, because the credential bundle handed to CI had drifted from the copy on disk — two copies of the same list, edited independently, which is its own recurring lesson.&lt;/p&gt;

&lt;p&gt;Every one of those is a five-minute fix once you can &lt;em&gt;see&lt;/em&gt; it. They had been invisible for over a week, and the instrument that was supposed to show them was reporting green the whole time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fixing it
&lt;/h2&gt;

&lt;p&gt;Any one of these is sufficient. They are not equivalent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Name the shell.&lt;/strong&gt; The smallest diff, and it fixes every pipeline in the step at once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run connector smoke tests&lt;/span&gt;
  &lt;span class="na"&gt;shell&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;bash&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;pytest tests/test_connector_smoke/ -v --tb=short -rs | tee /tmp/smoke.log&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. Set the flag explicitly.&lt;/strong&gt; More obvious to a reader, and it survives someone deleting &lt;code&gt;shell: bash&lt;/code&gt; as redundant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
  &lt;span class="s"&gt;set -o pipefail&lt;/span&gt;
  &lt;span class="s"&gt;pytest ... | tee /tmp/smoke.log&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3. Capture the status you care about.&lt;/strong&gt; Verbose, but it is the only form that lets you act on &lt;code&gt;pytest&lt;/code&gt;'s code specifically — useful when a tool distinguishes "tests failed" (&lt;code&gt;1&lt;/code&gt;) from "the run was misconfigured" (&lt;code&gt;2&lt;/code&gt;, &lt;code&gt;4&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pytest ... | &lt;span class="nb"&gt;tee&lt;/span&gt; /tmp/smoke.log
&lt;span class="nv"&gt;rc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;PIPESTATUS&lt;/span&gt;&lt;span class="p"&gt;[0]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;4. Stop piping.&lt;/strong&gt; Write the log to a file and &lt;code&gt;cat&lt;/code&gt; it afterwards, or emit machine-readable output and read that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pytest ... &lt;span class="nt"&gt;--junitxml&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/tmp/smoke.xml | &lt;span class="nb"&gt;tee&lt;/span&gt; /tmp/smoke.log
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the &lt;code&gt;grep&lt;/code&gt;, in every case. It still catches the case it was written for. It just must not be the &lt;em&gt;only&lt;/em&gt; thing that can fail the step.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not trust the green afterwards
&lt;/h2&gt;

&lt;p&gt;This is the step that gets skipped, and it is the one that matters.&lt;/p&gt;

&lt;p&gt;Once you have applied the fix, re-run the job &lt;strong&gt;unchanged, against the still-broken system&lt;/strong&gt;, and confirm it goes &lt;strong&gt;red&lt;/strong&gt;. If it goes green, your fix did not land — a check that has never been observed failing has never been shown able to fail. Our own rule, from a file of rules that each cost us an incident, is blunt about it: &lt;em&gt;a passing check is not evidence until you have seen it fail.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We have paid for that rule more than once. A monthly database restore drill asserted that a table of seed rows survived the restore — while &lt;code&gt;pg_dump&lt;/code&gt; writes tables alphabetically and puts &lt;code&gt;users&lt;/code&gt; at the very end of the file, which makes it the first thing a truncation destroys and the last thing that check would notice. It printed &lt;code&gt;PASS&lt;/code&gt; beside an empty user table. Alert rules that were structurally unable to fire. A test suite that mocked the very unit under test. In &lt;a href="https://datanika.io/blog/green-tests-broken-connectors/" rel="noopener noreferrer"&gt;an earlier post&lt;/a&gt; we wrote up 2,300 passing tests sitting beside a CSV import that loaded exactly one row.&lt;/p&gt;

&lt;p&gt;The through-line is one question, and it is worth asking of any signal before you rely on it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Would this look different if the thing it watches had failed?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you cannot answer that from the code, the green tells you nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Audit your own workflows in one command
&lt;/h2&gt;

&lt;p&gt;If you pipe test output anywhere in CI — to &lt;code&gt;tee&lt;/code&gt;, &lt;code&gt;grep&lt;/code&gt;, &lt;code&gt;head&lt;/code&gt;, &lt;code&gt;jq&lt;/code&gt;, a log shipper — check whether your exit code survives:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;f &lt;span class="k"&gt;in&lt;/span&gt; .github/workflows/&lt;span class="k"&gt;*&lt;/span&gt;.yml&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="s1"&gt;'pipefail'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="k"&gt;continue
  &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="s1"&gt;'shell: bash'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="k"&gt;continue
  &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-Hn&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'\|\s*(tee|grep|head|tail|jq|awk|sed|sort|uniq)'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every line it prints is a pipeline whose left-hand exit code is being discarded. Some of those are deliberate. The ones that are not are the ones running your tests.&lt;/p&gt;

&lt;p&gt;Two things worth knowing while you read the output:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;set -e&lt;/code&gt; does not save you.&lt;/strong&gt; &lt;code&gt;-e&lt;/code&gt; aborts on a failing &lt;em&gt;command&lt;/em&gt;; a pipeline whose last command succeeded has not failed, so there is nothing for &lt;code&gt;-e&lt;/code&gt; to abort on. Having &lt;code&gt;-e&lt;/code&gt; is what makes this feel safe when it is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Windows runners differ.&lt;/strong&gt; The default shell there is PowerShell, which has its own rules for &lt;code&gt;$LASTEXITCODE&lt;/code&gt; and native-command failure. A fix applied to a Linux matrix leg does not necessarily apply to a Windows one.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Alerting on the pipelines you actually run
&lt;/h2&gt;

&lt;p&gt;The uncomfortable part of this story is not that a workflow was misconfigured. It is that we had an alert wired up, pointed at the right workflow, with the right condition — and it was silent for eight nights because the condition it tested was never reached. The alert was fine. The thing it was reading was wrong.&lt;/p&gt;

&lt;p&gt;That generalises past CI. If you run scheduled data pipelines, the same question applies to every notification you have configured: does your failure alert read the pipeline's real outcome, or something downstream of a step that always succeeds?&lt;/p&gt;

&lt;p&gt;Datanika records the outcome of every run in its own ledger rather than inferring it from a log line, and failure notifications fire off that record. If you want them in Slack, &lt;a href="https://datanika.io/blog/slack-alerts-pipeline-failures/" rel="noopener noreferrer"&gt;that setup takes about two minutes&lt;/a&gt;. If you drive pipelines from CI, the &lt;a href="https://datanika.io/docs/api/" rel="noopener noreferrer"&gt;REST API&lt;/a&gt; hands back a run id and a status you can poll and assert on — which is a better thing to gate a deploy on than an exit code you have to hope survived a pipe.&lt;/p&gt;

&lt;p&gt;The platform is open source and self-hostable. &lt;a href="https://app.datanika.io" rel="noopener noreferrer"&gt;Start here&lt;/a&gt;, or &lt;a href="https://datanika.io/docs/" rel="noopener noreferrer"&gt;read the docs&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ci</category>
      <category>githubactions</category>
      <category>testing</category>
      <category>bash</category>
    </item>
    <item>
      <title>Triggering Data Pipelines from CI/CD via the REST API</title>
      <dc:creator>Evgenii Timofeev</dc:creator>
      <pubDate>Mon, 07 Sep 2026 12:44:43 +0000</pubDate>
      <link>https://dev.to/eu_ti_f127c5b5d7535b7174f/triggering-data-pipelines-from-cicd-via-the-rest-api-5027</link>
      <guid>https://dev.to/eu_ti_f127c5b5d7535b7174f/triggering-data-pipelines-from-cicd-via-the-rest-api-5027</guid>
      <description>&lt;p&gt;The most common thing people want from a data platform's API is boring: &lt;strong&gt;run this pipeline, tell me if it worked.&lt;/strong&gt; Usually right after a deploy, a dbt seed change, or a nightly job that has to finish before something else starts.&lt;/p&gt;

&lt;p&gt;Here's how to do that with Datanika — and, more usefully, the one place this is easy to get wrong in a way your CI won't notice. We got it wrong ourselves; the fix is at the end of that section.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;--fail-with-body&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"https://app.datanika.io/api/v1/pipelines/1/run?wait=true&amp;amp;timeout=300"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$DATANIKA_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One call. Blocks until the pipeline finishes, returns the run as JSON, and &lt;strong&gt;exits non-zero if the pipeline failed&lt;/strong&gt;. In a &lt;code&gt;set -e&lt;/code&gt; CI step, that is the whole integration.&lt;/p&gt;

&lt;p&gt;The reason that last clause is worth a sentence is the rest of this section.&lt;/p&gt;

&lt;h2&gt;
  
  
  The status code carries the run's outcome
&lt;/h2&gt;

&lt;p&gt;With &lt;code&gt;?wait=true&lt;/code&gt;, the endpoint polls the run until it reaches a terminal state, then answers with a status code that describes &lt;strong&gt;the run&lt;/strong&gt;, not just the request:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The run finished successfully&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;200&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Still pending or running when your timeout expired&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;408&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal, but not successful (&lt;code&gt;failed&lt;/code&gt;, &lt;code&gt;cancelled&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;422&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The body is the serialized run in all three cases, so &lt;code&gt;status&lt;/code&gt; and &lt;code&gt;error_message&lt;/code&gt; are still there when you want detail:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;43&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"target_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pipeline"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"target_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"failed"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"started_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-08-30T14:00:02Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"finished_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-08-30T14:00:44Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"rows_loaded"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"error_message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"relation &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;public.orders&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt; does not exist"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"created_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-08-30T14:00:00Z"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That response is a &lt;strong&gt;422&lt;/strong&gt;. &lt;code&gt;curl --fail&lt;/code&gt; trips on it, &lt;code&gt;raise_for_status()&lt;/code&gt; raises on it, and your CI job goes red — which is what you wanted when you asked the API to wait.&lt;/p&gt;

&lt;h3&gt;
  
  
  The trap this replaced, because you will meet it elsewhere
&lt;/h3&gt;

&lt;p&gt;Until August 2026, that same failed run came back as &lt;strong&gt;HTTP 200&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The reasoning was defensible and wrong. Your &lt;em&gt;request&lt;/em&gt; succeeded — the server did its job, found the run, waited, serialized it, and handed it to you. The run is what failed, and the body says so plainly in &lt;code&gt;status&lt;/code&gt;. Textbook HTTP.&lt;/p&gt;

&lt;p&gt;The problem is that nothing in CI reads the body:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# The reflex every CI script has&lt;/span&gt;
curl &lt;span class="nt"&gt;--fail&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;".../pipelines/1/run?wait=true"&lt;/span&gt; &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--fail&lt;/code&gt; trips on 4xx and 5xx. A failed pipeline was a 200. &lt;strong&gt;Exit code 0.&lt;/strong&gt; The job goes green, the dashboard downstream is stale, and nobody finds out until someone notices the numbers are yesterday's. A step that reports success on a failed load is worse than no step at all, because it converts a loud failure into a silent one.&lt;/p&gt;

&lt;p&gt;What settled it was noticing the endpoint had &lt;strong&gt;already&lt;/strong&gt; decided the question. It returned &lt;strong&gt;408&lt;/strong&gt; when the run was still going at the timeout — and that isn't a transport failure either; the request was served perfectly. So &lt;code&gt;200 == failed&lt;/code&gt; wasn't a competing philosophy, it was an inconsistency with the endpoint's own behaviour. &lt;code&gt;?wait=true&lt;/code&gt; is the caller explicitly opting into &lt;em&gt;"block until you know the outcome."&lt;/em&gt; If the outcome doesn't reach the status line, the option is half-built.&lt;/p&gt;

&lt;p&gt;So if you are integrating some &lt;em&gt;other&lt;/em&gt; pipeline API and it offers a "wait" mode: &lt;strong&gt;check what a failed run returns before you trust &lt;code&gt;--fail&lt;/code&gt;.&lt;/strong&gt; Trigger something you know is broken and look at the exit code. It is a two-minute experiment that this post exists because we ran late.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why 422 and not 500
&lt;/h3&gt;

&lt;p&gt;A 5xx means &lt;em&gt;"our API broke, retry the request."&lt;/em&gt; Retrying this request would start a &lt;strong&gt;second pipeline run&lt;/strong&gt; — a second extract, a second load, a second set of rows. A failed pipeline is not a transport failure and must not be retried like one.&lt;/p&gt;

&lt;p&gt;422 says the opposite: the request was fine, and the thing you asked about did not succeed. The failure is in your pipeline — your credentials, your SQL, your source — not in our server. Retrying blindly is exactly the wrong move, and the status code should say so.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;cancelled&lt;/code&gt; is a 422 as well. The check is &lt;em&gt;"not success"&lt;/em&gt; rather than a list of failure names, so a terminal status added later cannot quietly rejoin the success branch.&lt;/p&gt;

&lt;h3&gt;
  
  
  Keep the body when you fail
&lt;/h3&gt;

&lt;p&gt;Plain &lt;code&gt;curl --fail&lt;/code&gt; discards the response body on an HTTP error, which throws away &lt;code&gt;error_message&lt;/code&gt; — the one thing you want in the log.&lt;/p&gt;

&lt;p&gt;Use &lt;strong&gt;&lt;code&gt;--fail-with-body&lt;/code&gt;&lt;/strong&gt; (curl 7.76+, so every current GitHub runner) to get the non-zero exit &lt;em&gt;and&lt;/em&gt; the body. If you're on something older, capture the status code explicitly, as the full workflow below does.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three endpoints
&lt;/h2&gt;

&lt;p&gt;Runs are triggered on the &lt;strong&gt;resource&lt;/strong&gt;, not on a runs collection. There is no &lt;code&gt;POST /api/v1/runs&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you're running&lt;/th&gt;
&lt;th&gt;Endpoint&lt;/th&gt;
&lt;th&gt;Key scope needed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A pipeline&lt;/td&gt;
&lt;td&gt;&lt;code&gt;POST /api/v1/pipelines/{id}/run&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;pipelines:write&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A file upload&lt;/td&gt;
&lt;td&gt;&lt;code&gt;POST /api/v1/uploads/{id}/run&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;uploads:write&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A dbt transformation&lt;/td&gt;
&lt;td&gt;&lt;code&gt;POST /api/v1/transformations/{id}/run&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;transformations:write&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All three behave identically with respect to everything below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two modes: fire-and-forget, or wait
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Without &lt;code&gt;wait&lt;/code&gt;&lt;/strong&gt; you get an immediate &lt;code&gt;202 Accepted&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"run_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;43&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pending"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use this when CI's job is to &lt;em&gt;kick off&lt;/em&gt; work — a nightly ingest that takes 40 minutes and nothing downstream is blocking on it. A &lt;code&gt;202&lt;/code&gt; is about dispatch only; it says nothing about the outcome, and there is nothing to check with &lt;code&gt;--fail&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With &lt;code&gt;?wait=true&lt;/code&gt;&lt;/strong&gt; the request blocks until the run is terminal, then answers 200 / 408 / 422 as above.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;timeout&lt;/code&gt; is in seconds, defaults to &lt;strong&gt;120&lt;/strong&gt;, and is clamped to &lt;strong&gt;1–300&lt;/strong&gt;. Passing &lt;code&gt;timeout=3600&lt;/code&gt; doesn't get you an hour — you get 300 seconds, then a 408. Status is polled every 2 seconds, and waiting doesn't occupy a worker, so a waiting request costs you nothing but the open connection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A 408 is not a failure.&lt;/strong&gt; The run is still going; you just stopped waiting. The body carries &lt;code&gt;"timed_out": true&lt;/code&gt;, and the run's own &lt;code&gt;status&lt;/code&gt; is still &lt;code&gt;pending&lt;/code&gt; or &lt;code&gt;running&lt;/code&gt;. That distinction matters because the remedy is different — a 422 means fix your pipeline, a 408 means wait longer or stop blocking on it. Treating them the same is how a slow Tuesday becomes a red build.&lt;/p&gt;

&lt;h2&gt;
  
  
  Runs longer than five minutes
&lt;/h2&gt;

&lt;p&gt;Because of that 300-second ceiling, anything longer needs the async shape: trigger, then poll.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;run_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;--fail-with-body&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"https://app.datanika.io/api/v1/pipelines/1/run"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$DATANIKA_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.run_id'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="nv"&gt;deadline&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;$((&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%s&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="m"&gt;3600&lt;/span&gt; &lt;span class="k"&gt;))&lt;/span&gt;
&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%s&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-lt&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$deadline&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;&lt;span class="nv"&gt;run&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="s2"&gt;"https://app.datanika.io/api/v1/runs/&lt;/span&gt;&lt;span class="nv"&gt;$run_id&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$DATANIKA_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
  &lt;span class="nv"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.status'&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$run&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$status&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="k"&gt;in
    &lt;/span&gt;success&lt;span class="p"&gt;)&lt;/span&gt;             &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"done"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;0 &lt;span class="p"&gt;;;&lt;/span&gt;
    failed|cancelled&lt;span class="p"&gt;)&lt;/span&gt;    jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.error_message // .status'&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$run&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1 &lt;span class="p"&gt;;;&lt;/span&gt;
  &lt;span class="k"&gt;esac&lt;/span&gt;
  &lt;span class="nb"&gt;sleep &lt;/span&gt;15
&lt;span class="k"&gt;done

&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Still running after 1h — check run &lt;/span&gt;&lt;span class="nv"&gt;$run_id&lt;/span&gt;&lt;span class="s2"&gt; in the app"&lt;/span&gt;
&lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;GET /api/v1/runs/{id}&lt;/code&gt; is a plain read: it returns 200 with the run whatever its status, because here you &lt;em&gt;are&lt;/em&gt; asking about the row rather than asking "did it work?". The &lt;code&gt;.status&lt;/code&gt; field is the answer in this shape.&lt;/p&gt;

&lt;p&gt;Note the &lt;code&gt;case&lt;/code&gt; covers &lt;strong&gt;every&lt;/strong&gt; terminal status, not just &lt;code&gt;success&lt;/code&gt;. A loop that only watches for &lt;code&gt;success&lt;/code&gt; runs until your CI timeout and then reports the wrong cause.&lt;/p&gt;

&lt;p&gt;If you want the run's output while debugging, &lt;code&gt;GET /api/v1/runs/{id}/logs&lt;/code&gt; returns &lt;code&gt;{"run_id": …, "logs": "…"}&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Idempotent retries: the header that stops double-runs
&lt;/h2&gt;

&lt;p&gt;CI reruns. Someone clicks "Re-run failed jobs", a runner gets evicted mid-step, a network blip makes &lt;code&gt;curl&lt;/code&gt; retry. Any of those can fire the same trigger twice — and by default, twice means two runs, two loads, and two sets of rows.&lt;/p&gt;

&lt;p&gt;Every &lt;code&gt;POST&lt;/code&gt; endpoint accepts an optional &lt;strong&gt;&lt;code&gt;Idempotency-Key&lt;/code&gt;&lt;/strong&gt; header. Replay the same key and you get the original response back instead of a second run. Keys are cached for &lt;strong&gt;24 hours&lt;/strong&gt;, and it's opt-in — no header, no deduplication.&lt;/p&gt;

&lt;p&gt;The natural key in GitHub Actions is the run identity itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;--fail-with-body&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"https://app.datanika.io/api/v1/pipelines/1/run?wait=true&amp;amp;timeout=300"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$DATANIKA_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Idempotency-Key: gha-&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;GITHUB_RUN_ID&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;-&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;GITHUB_RUN_ATTEMPT&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Include &lt;code&gt;GITHUB_RUN_ATTEMPT&lt;/code&gt; if a manual re-run &lt;em&gt;should&lt;/em&gt; start a fresh pipeline, and leave it out if it shouldn't. That's a real decision, not a formality — decide it deliberately.&lt;/p&gt;

&lt;h2&gt;
  
  
  A complete GitHub Actions job
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Refresh analytics&lt;/span&gt;

&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;push&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;branches&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;main&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dbt/**"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;workflow_dispatch&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;refresh&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Trigger the Datanika pipeline&lt;/span&gt;
        &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;DATANIKA_API_KEY&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.DATANIKA_API_KEY }}&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;set -euo pipefail&lt;/span&gt;

          &lt;span class="s"&gt;response=$(curl -sS -X POST \&lt;/span&gt;
            &lt;span class="s"&gt;"https://app.datanika.io/api/v1/pipelines/1/run?wait=true&amp;amp;timeout=300" \&lt;/span&gt;
            &lt;span class="s"&gt;-H "Authorization: Bearer $DATANIKA_API_KEY" \&lt;/span&gt;
            &lt;span class="s"&gt;-H "Idempotency-Key: gha-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}" \&lt;/span&gt;
            &lt;span class="s"&gt;-w "\n%{http_code}")&lt;/span&gt;

          &lt;span class="s"&gt;code=$(tail -n1 &amp;lt;&amp;lt;&amp;lt;"$response")&lt;/span&gt;
          &lt;span class="s"&gt;body=$(sed '$d' &amp;lt;&amp;lt;&amp;lt;"$response")&lt;/span&gt;

          &lt;span class="s"&gt;case "$code" in&lt;/span&gt;
            &lt;span class="s"&gt;200) echo "Loaded $(jq -r '.rows_loaded' &amp;lt;&amp;lt;&amp;lt;"$body") rows" ;;&lt;/span&gt;
            &lt;span class="s"&gt;408) echo "::warning::Still running after 300s — run $(jq -r '.id' &amp;lt;&amp;lt;&amp;lt;"$body")"&lt;/span&gt;
                 &lt;span class="s"&gt;exit 1 ;;&lt;/span&gt;
            &lt;span class="s"&gt;422) echo "::error::Pipeline run $(jq -r '.status' &amp;lt;&amp;lt;&amp;lt;"$body") — $(jq -r '.error_message // "no message"' &amp;lt;&amp;lt;&amp;lt;"$body")"&lt;/span&gt;
                 &lt;span class="s"&gt;exit 1 ;;&lt;/span&gt;
            &lt;span class="s"&gt;*)   echo "::error::API returned $code"; echo "$body"; exit 1 ;;&lt;/span&gt;
          &lt;span class="s"&gt;esac&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four branches, because there are four genuinely different situations: it worked, it's still going, your pipeline broke, or our API did. &lt;code&gt;--fail-with-body&lt;/code&gt; collapses the middle two into one non-zero exit, which is fine when you only need pass/fail — but the messages above are what you'll want at 2am, and the 408 branch is the one people most often want to handle differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cancelling a run
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Update, 15 September 2026.&lt;/strong&gt; One clause in this section is no longer true. A run cancelled while it is running now keeps the status &lt;code&gt;cancelled&lt;/code&gt; when its task finishes; it is no longer overwritten back to &lt;code&gt;success&lt;/code&gt;. Everything else here still holds: cancelling does not stop the worker, the extract and the load run to the end, and the run is billed for everything it processed. The original text follows as it was published.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If your workflow is cancelled, the pipeline it started is not. A &lt;code&gt;202&lt;/code&gt; handed the work to a background worker, and CI walking away doesn't reach it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And the API cannot currently clean it up for you — plan around that rather than around the endpoint.&lt;/strong&gt; &lt;code&gt;POST /api/v1/runs/{id}/cancel&lt;/code&gt; exists and will return &lt;code&gt;200&lt;/code&gt; with &lt;code&gt;"status": "cancelled"&lt;/code&gt; — but that call only writes the status onto the run row. It does not stop the worker. The extract keeps reading, the load keeps writing, and when the task finishes it overwrites the status back to &lt;code&gt;success&lt;/code&gt;. Tracked as &lt;a href="https://github.com/datanika-io/datanika-core/issues/657" rel="noopener noreferrer"&gt;core#657&lt;/a&gt;, along with the reason there is no cancel button in the app either: shipping a control onto that behaviour would spread the wrong impression to every user instead of only to API callers.&lt;/p&gt;

&lt;p&gt;So don't wire a cleanup step and believe it. Until cancellation actually stops work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bound the blast radius instead of the run.&lt;/strong&gt; Smaller, more frequent pipelines beat one long job you might want to kill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assume anything you triggered will finish.&lt;/strong&gt; Check the app rather than your CI log for what a cancelled workflow left behind.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On usage-based plans, a run you "cancelled" keeps metering&lt;/strong&gt; until it completes on its own. That's the practical reason this is a limitation worth stating rather than a detail.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The one accurate part of the old advice: the endpoint returns &lt;strong&gt;409&lt;/strong&gt; with &lt;code&gt;"not_cancellable"&lt;/code&gt; when the run has already finished, which is a perfectly normal outcome and not something to treat as an error.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give CI its own key, scoped down
&lt;/h2&gt;

&lt;p&gt;API keys are created in &lt;strong&gt;Settings → API Keys&lt;/strong&gt; in the app, carry the &lt;code&gt;etf_&lt;/code&gt; prefix, and are shown &lt;strong&gt;once&lt;/strong&gt;. They're hashed with SHA-256 before storage, so a lost key can't be recovered — you create a new one and revoke the old. Full details on &lt;a href="https://datanika.io/api/keys/" rel="noopener noreferrer"&gt;the API keys page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For a CI key, set the scopes explicitly rather than leaving them empty (empty means full access):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;pipelines:write&lt;/code&gt; — to trigger&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;runs:read&lt;/code&gt; — to poll status and read logs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's a key that can start and observe one kind of work and cannot delete a connection, read your credentials, or create a schedule. If it leaks into a build log, the blast radius is a pipeline someone can already trigger from the UI.&lt;/p&gt;

&lt;p&gt;Set an expiry on it too, and rotate it on a calendar rather than after an incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rate limits
&lt;/h2&gt;

&lt;p&gt;Each key is rate-limited independently, per plan, and exceeding it returns &lt;strong&gt;429&lt;/strong&gt; with a &lt;code&gt;Retry-After&lt;/code&gt; header. Current per-plan limits are in &lt;a href="https://datanika.io/api/reference/#rate-limits" rel="noopener noreferrer"&gt;the API reference&lt;/a&gt; — a fan-out matrix build that triggers one pipeline per shard is the realistic way to hit them, so back off on 429 rather than retrying immediately.&lt;/p&gt;

&lt;p&gt;The polling loop above is well inside every tier: at &lt;code&gt;sleep 15&lt;/code&gt; it costs 4 requests a minute, and &lt;code&gt;?wait=true&lt;/code&gt; polls &lt;strong&gt;server-side&lt;/strong&gt;, so a 300-second wait is one request against your budget, not 150.&lt;/p&gt;

&lt;p&gt;Note that 429 is a genuine transport-level "try again", unlike the 422 above. It's the one 4xx here you &lt;em&gt;should&lt;/em&gt; retry.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is easier here than in a three-tool stack
&lt;/h2&gt;

&lt;p&gt;The reason this post is short is architectural. In a Fivetran + dbt Cloud + Airflow stack, "run the pipeline and tell me if it worked" is three APIs, three auth schemes, three status vocabularies, and a decision about which failure counts. Here, extract, load, and transform are the same run object with one &lt;code&gt;status&lt;/code&gt; field, so CI asks one question once.&lt;/p&gt;

&lt;p&gt;That argument, with numbers attached, is in &lt;a href="https://datanika.io/blog/datanika-vs-modern-data-stack/" rel="noopener noreferrer"&gt;Datanika vs the Modern Data Stack&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this post does not cover
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Scheduling.&lt;/strong&gt; If you want a pipeline to run nightly, use a &lt;a href="https://datanika.io/docs/scheduling/" rel="noopener noreferrer"&gt;schedule&lt;/a&gt; instead of a cron-triggered CI job. Cron in CI gives you the worst of both — an extra dependency and no visibility in the app.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Webhooks back into CI.&lt;/strong&gt; There is no outbound "run finished" callback into a workflow today; poll, or use a &lt;a href="https://datanika.io/blog/slack-alerts-pipeline-failures/" rel="noopener noreferrer"&gt;notification channel&lt;/a&gt; to tell a human.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every endpoint.&lt;/strong&gt; The &lt;a href="https://datanika.io/api/reference/" rel="noopener noreferrer"&gt;API reference&lt;/a&gt; has the full surface, including connections, schedules, and bulk import.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Every status code, header, scope name and default here was read out of the live OpenAPI document and the shipped route handlers, not out of our own docs. That matters more than usual for this post: the 200-on-failure behaviour in "the trap this replaced" was **found while writing the first draft of it&lt;/em&gt;&lt;em&gt;, filed as &lt;a href="https://github.com/datanika-io/datanika-core/issues/663" rel="noopener noreferrer"&gt;core#663&lt;/a&gt;, and fixed before this went out. Writing the tutorial is what surfaced the bug — so if you find a discrepancy, &lt;a href="https://github.com/datanika-io/datanika-landing/issues" rel="noopener noreferrer"&gt;open an issue&lt;/a&gt; and we'll fix the post, or the API.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>restapi</category>
      <category>cicd</category>
      <category>githubactions</category>
      <category>automation</category>
    </item>
  </channel>
</rss>
