<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gowtham Potureddi</title>
    <description>The latest articles on DEV Community by Gowtham Potureddi (@gowthampotureddi).</description>
    <link>https://dev.to/gowthampotureddi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3874592%2Fb901f929-0a60-4dd2-9dac-22ce22291bdc.png</url>
      <title>DEV Community: Gowtham Potureddi</title>
      <link>https://dev.to/gowthampotureddi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gowthampotureddi"/>
    <language>en</language>
    <item>
      <title>Prophecy: Visual, Git-Backed Low-Code Spark &amp; SQL Pipelines for the Enterprise</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Sun, 27 Sep 2026 17:47:57 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/prophecy-visual-git-backed-low-code-spark-sql-pipelines-for-the-enterprise-ik8</link>
      <guid>https://dev.to/gowthampotureddi/prophecy-visual-git-backed-low-code-spark-sql-pipelines-for-the-enterprise-ik8</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;Prophecy low-code data engineering&lt;/code&gt;&lt;/strong&gt; is a visual, drag-and-drop pipeline builder that does something most low-code tools do not: instead of running your work inside a proprietary black box, it &lt;strong&gt;compiles the canvas to open-source code&lt;/strong&gt; — readable PySpark or Scala Spark for data pipelines, or dbt-style SQL models for warehouse pipelines — and commits that code to your own Git repository. You wire a few boxes together on a graph, and Prophecy generates the exact Spark job a senior engineer would have hand-written, tests and all.&lt;/p&gt;

&lt;p&gt;That is a genuinely different shape from the two options enterprise data teams reached for before it: a legacy visual ETL suite (Informatica, Alteryx, Ab Initio, SSIS) whose logic is trapped in an opaque metadata format only that vendor can execute, or a purely hand-coded Spark stack that only your most senior engineers can safely touch. Prophecy sits in the middle — visual enough that analysts and junior engineers are productive, but code-first enough that the output is standard, version-controlled, runnable-without-the-vendor Spark and SQL. This guide walks through the five ideas an interviewer or an architecture-review board will actually probe — the gem-and-pipeline model, the visual↔code round-trip on Git, execution on Spark and Databricks via Fabrics, and reusable subgraphs with tests and column-level lineage — and pairs each with a Solution-Tail interview answer: code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl0rufqgaf16hzbjsxfdd.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl0rufqgaf16hzbjsxfdd.jpeg" alt="PipeCode blog header for Prophecy — bold white headline 'Prophecy: Visual → Code' with subtitle 'Git-Backed Low-Code Spark &amp;amp; SQL Pipelines' and a stylised visual-canvas-compiles-to-code scene on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/spark-sql" rel="noopener noreferrer"&gt;Spark SQL practice library →&lt;/a&gt;, rehearse the extract-and-load shapes on the &lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;ETL practice set →&lt;/a&gt;, and sharpen the transform logic Prophecy generates on the &lt;a href="https://pipecode.ai/explore/practice/topic/data-transformation" rel="noopener noreferrer"&gt;data-transformation practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why Prophecy changes enterprise pipeline building in 2026&lt;/li&gt;
&lt;li&gt;Gems, pipelines &amp;amp; the visual canvas&lt;/li&gt;
&lt;li&gt;The visual ↔ code round-trip on Git&lt;/li&gt;
&lt;li&gt;Running on Spark &amp;amp; Databricks with Fabrics&lt;/li&gt;
&lt;li&gt;Reusable subgraphs, tests &amp;amp; column-level lineage&lt;/li&gt;
&lt;li&gt;Cheat sheet — Prophecy recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why Prophecy changes enterprise pipeline building in 2026
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Prophecy is a code generator wearing a visual UI — that one fact decides where it fits
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;Prophecy is not a runtime you send data through; it is a visual editor that emits ordinary Spark and SQL code, which then runs on your own engine&lt;/strong&gt;. Everything that makes Prophecy defensible to a platform team follows from that. There is no proprietary execution engine to license per row, no metadata format only the vendor can read, and no "export to a dead PDF of a diagram" migration cliff — the artifact Prophecy produces is a Git repo full of PySpark files or SQL models you could keep running if Prophecy vanished tomorrow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The low-code-without-lock-in split — what Prophecy does and deliberately does not do.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Authoring.&lt;/strong&gt; Prophecy gives you a drag-and-drop canvas of &lt;strong&gt;gems&lt;/strong&gt; (transform nodes) that snap into a pipeline graph, with per-gem data previews and a schema that flows through the graph as you build.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code generation.&lt;/strong&gt; Every visual pipeline compiles to standard code: &lt;strong&gt;PySpark or Scala&lt;/strong&gt; for Spark projects, &lt;strong&gt;dbt-style SQL models&lt;/strong&gt; for SQL projects. That code is the source of truth and lives in Git.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Execution is out of scope on purpose.&lt;/strong&gt; Prophecy does not run your data on its own cluster; the generated Spark runs on &lt;strong&gt;Databricks (or any Spark)&lt;/strong&gt;, and the generated SQL runs on your &lt;strong&gt;warehouse via dbt&lt;/strong&gt;. Prophecy stays a design-and-generate layer, which is exactly why the output has no lock-in.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Where Prophecy sits against the alternatives.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;vs hand-written Spark.&lt;/strong&gt; A raw PySpark codebase is maximally flexible but gates every change behind a senior engineer and a code review. Prophecy lets an analyst build the same pipeline visually while still producing the same reviewable PySpark — you widen the contributor pool without lowering the artifact quality.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vs legacy Informatica / Alteryx / Ab Initio.&lt;/strong&gt; Those visual suites also let non-experts build pipelines, but the logic is trapped in a vendor metadata format and runs only on the vendor engine. Prophecy's visual output is open-source Spark/SQL in your repo, so there is no proprietary runtime to keep paying for and no black box to reverse-engineer during an audit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vs Fivetran / Airbyte.&lt;/strong&gt; Managed connectors solve &lt;strong&gt;extract-load&lt;/strong&gt; for sources that already have a connector. Prophecy solves &lt;strong&gt;transform&lt;/strong&gt; — the joins, aggregations, and business logic downstream of ingestion — and generates the Spark/SQL that expresses it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you say &lt;strong&gt;"Prophecy compiles the canvas to code, it is not a runtime"&lt;/strong&gt; in the first sentence? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you frame the value as &lt;strong&gt;"low-code authoring, code-first artifact"&lt;/strong&gt; rather than "no-code magic"? — required framing.&lt;/li&gt;
&lt;li&gt;Do you reach for Prophecy when the goal is &lt;strong&gt;"migrate Informatica to Spark without hand-rewriting 4,000 mappings"&lt;/strong&gt; or &lt;strong&gt;"let analysts contribute reviewable Spark"&lt;/strong&gt;, not as a Fivetran replacement? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you mention &lt;strong&gt;Git, tests, and column-level lineage as built-ins&lt;/strong&gt;, not afterthoughts? — the enterprise point.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — three boxes that become a real Spark job
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The canonical Prophecy "hello world" is a three-gem pipeline — a Source that reads a table, a Reformat that derives a column, and a Target that writes the result. It looks like a toy flowchart, and that is the point: the same three boxes that transform a five-row table generate the same PySpark structure that a production job would, because Prophecy always emits one function per gem and a &lt;code&gt;main&lt;/code&gt; that wires them together. No notebook, no boilerplate &lt;code&gt;SparkSession&lt;/code&gt; glue that you maintain by hand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Read an &lt;code&gt;orders&lt;/code&gt; table, add a &lt;code&gt;total = qty * price&lt;/code&gt; column, and write to &lt;code&gt;orders_enriched&lt;/code&gt;. Show the shape of the Spark code Prophecy generates from the three gems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;qty&lt;/th&gt;
&lt;th&gt;price&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;10.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;4.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;functions&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;Source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;table&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sales.orders&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;Reformat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;total&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;qty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;Target&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;overwrite&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;saveAsTable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sales.orders_enriched&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;pipeline&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;df_src&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;df_ref&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Reformat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;df_src&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nc"&gt;Target&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;df_ref&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The &lt;strong&gt;Source&lt;/strong&gt; gem becomes a function that returns a &lt;code&gt;DataFrame&lt;/code&gt; from &lt;code&gt;sales.orders&lt;/code&gt;; it declares &lt;em&gt;where data comes from&lt;/em&gt; and nothing else. The &lt;strong&gt;Reformat&lt;/strong&gt; gem becomes a pure &lt;code&gt;DataFrame → DataFrame&lt;/code&gt; function that adds the &lt;code&gt;total&lt;/code&gt; column with &lt;code&gt;withColumn&lt;/code&gt;. The &lt;strong&gt;Target&lt;/strong&gt; gem becomes a sink function that writes the table. The &lt;code&gt;pipeline&lt;/code&gt; driver wires them in dependency order — exactly the DAG you drew on the canvas — so the visual graph and the code are the same object viewed two ways.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;qty&lt;/th&gt;
&lt;th&gt;price&lt;/th&gt;
&lt;th&gt;total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;10.0&lt;/td&gt;
&lt;td&gt;20.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;4.0&lt;/td&gt;
&lt;td&gt;12.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If you can draw the pipeline as boxes and arrows, Prophecy can generate it as functions and a driver — the toy example and the 200-gem production pipeline differ only in how many gems there are, never in the code shape.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Gems, pipelines &amp;amp; the visual canvas
&lt;/h2&gt;

&lt;h3&gt;
  
  
  A gem is one transform node, a pipeline is a DAG of gems, and each gem compiles to a function
&lt;/h3&gt;

&lt;p&gt;Prophecy has one authoring primitive you compose endlessly — the &lt;strong&gt;gem&lt;/strong&gt; — and an interviewer who asks "how is a Prophecy pipeline structured?" wants the gem-to-pipeline-to-code chain in order. Get this vocabulary crisp and the whole tool snaps into focus.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The core objects.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gem — one transform node.&lt;/strong&gt; A gem is a single box on the canvas that performs one operation. Built-in gems cover the standard Spark verbs: &lt;strong&gt;Source&lt;/strong&gt; and &lt;strong&gt;Target&lt;/strong&gt; (read/write), &lt;strong&gt;Reformat&lt;/strong&gt; (select/derive columns), &lt;strong&gt;Aggregate&lt;/strong&gt; (group-by), &lt;strong&gt;Join&lt;/strong&gt;, &lt;strong&gt;Filter&lt;/strong&gt;, &lt;strong&gt;OrderBy&lt;/strong&gt;, &lt;strong&gt;Deduplicate&lt;/strong&gt;, &lt;strong&gt;SetOperation&lt;/strong&gt; (union/intersect), and &lt;strong&gt;FlattenSchema&lt;/strong&gt; for nested data. Each gem carries its own configuration and its own output schema.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pipeline — a DAG of gems.&lt;/strong&gt; Wiring gem outputs to gem inputs builds a directed acyclic graph. Data flows as typed Spark &lt;code&gt;DataFrame&lt;/code&gt;s along the edges; the schema at each edge is known at design time, so Prophecy can validate column references before you ever run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Project — a Git-backed collection of pipelines.&lt;/strong&gt; A project bundles pipelines, subgraphs, tests, and generated code into one repository. It is the unit that maps to a Git repo.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What every gem gives you.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A generated function.&lt;/strong&gt; Each gem compiles to exactly one function in the project code — &lt;code&gt;Reformat(spark, in0)&lt;/code&gt;, &lt;code&gt;Aggregate(spark, in0)&lt;/code&gt;, and so on — so the code is as modular as the diagram.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A visual expression builder.&lt;/strong&gt; Inside a gem you write column expressions in SQL-like syntax (&lt;code&gt;qty * price&lt;/code&gt;, &lt;code&gt;upper(name)&lt;/code&gt;), and Prophecy translates them into the target dialect (Spark &lt;code&gt;functions&lt;/code&gt; calls for PySpark, raw SQL for SQL projects).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An interactive preview.&lt;/strong&gt; You can run the pipeline up to any gem and see a sample of the &lt;code&gt;DataFrame&lt;/code&gt; at that point, which is the low-code equivalent of a breakpoint.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why the gem is the powerful unit.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Because each gem is an isolated function, a pipeline is trivially testable gem-by-gem and readable diff-by-diff in a pull request.&lt;/li&gt;
&lt;li&gt;Because gems carry schema, Prophecy catches "column does not exist" at design time, not at 3 a.m. in production.&lt;/li&gt;
&lt;li&gt;Because the gem set is extensible, a platform team can ship &lt;strong&gt;custom gems&lt;/strong&gt; (via the Gem Builder) that encode company standards, and every analyst gets them as drag-and-drop boxes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm52ig205405wqar33opl.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm52ig205405wqar33opl.jpeg" alt="Iconographic Prophecy canvas diagram — source, reformat, aggregate, join and target gems wired into a pipeline DAG, each gem compiling to a function in the generated code, with a per-gem data-preview chip." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — a Filter gem and a Reformat gem in one pipeline
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; Real pipelines chain several gems. Here a two-gem transform keeps only paid orders (a Filter gem) and then standardizes the customer name (a Reformat gem). Each gem becomes its own function, and the pipeline driver composes them — the canvas order is the call order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Build a pipeline that filters &lt;code&gt;orders&lt;/code&gt; to &lt;code&gt;status = 'paid'&lt;/code&gt; and uppercases &lt;code&gt;customer&lt;/code&gt;, and show the two generated gem functions plus the driver.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; Four orders, two of them paid, returned by the Source gem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;functions&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;FilterPaid&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;paid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;StandardizeCustomer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;upper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;pipeline&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;src&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;table&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sales.orders&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;paid&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FilterPaid&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;clean&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;StandardizeCustomer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;paid&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;clean&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;overwrite&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;saveAsTable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sales.orders_clean&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The &lt;strong&gt;Filter&lt;/strong&gt; gem becomes &lt;code&gt;FilterPaid&lt;/code&gt;, a &lt;code&gt;DataFrame → DataFrame&lt;/code&gt; function applying one predicate. The &lt;strong&gt;Reformat&lt;/strong&gt; gem becomes &lt;code&gt;StandardizeCustomer&lt;/code&gt;, applying one &lt;code&gt;withColumn&lt;/code&gt;. The driver reads the source, threads it through the two gem functions in canvas order, and writes the sink. Because each gem is a separate function, a reviewer reading the pull request sees exactly two logical changes, cleanly named.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;ADA&lt;/td&gt;
&lt;td&gt;paid&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;LINUS&lt;/td&gt;
&lt;td&gt;paid&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; One gem = one function = one reviewable unit. Keep each gem doing a single verb and your generated code stays as clean as your diagram.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prophecy interview question on the gem/pipeline model
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; An interviewer describes a Reformat gem that should derive three columns — &lt;code&gt;total = qty * price&lt;/code&gt;, &lt;code&gt;year = year(order_ts)&lt;/code&gt;, and a normalized &lt;code&gt;email = lower(trim(email))&lt;/code&gt; — from an &lt;code&gt;orders&lt;/code&gt; DataFrame. Write the single Spark transform this gem generates, keeping it one projection so the physical plan stays flat.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using a single projection with column expressions
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;functions&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;Reformat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;order_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;qty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;qty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;alias&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;total&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;year&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;order_ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;alias&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;year&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))).&lt;/span&gt;&lt;span class="nf"&gt;alias&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;expression applied&lt;/th&gt;
&lt;th&gt;column produced&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;qty * price&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;total&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;&lt;code&gt;year(order_ts)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;year&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;lower(trim(email))&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;email&lt;/code&gt; (overwritten)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;one &lt;code&gt;select&lt;/code&gt; projection&lt;/td&gt;
&lt;td&gt;flat plan&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;A Reformat gem is a &lt;strong&gt;projection&lt;/strong&gt;: it maps input columns to output columns in a single &lt;code&gt;select&lt;/code&gt;, so all derivations land in one plan node.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;total&lt;/code&gt; multiplies two columns; &lt;code&gt;year&lt;/code&gt; extracts the year from a timestamp; &lt;code&gt;email&lt;/code&gt; is normalized by nesting &lt;code&gt;lower&lt;/code&gt; around &lt;code&gt;trim&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Because everything is one &lt;code&gt;select&lt;/code&gt;, Spark's Catalyst optimizer fuses the expressions into a single project stage — no extra shuffle or scan.&lt;/li&gt;
&lt;li&gt;Naming each derived column with &lt;code&gt;.alias(...)&lt;/code&gt; is what makes the gem's output schema explicit and lets the next gem reference &lt;code&gt;total&lt;/code&gt; and &lt;code&gt;year&lt;/code&gt; by name.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;qty&lt;/th&gt;
&lt;th&gt;price&lt;/th&gt;
&lt;th&gt;total&lt;/th&gt;
&lt;th&gt;year&lt;/th&gt;
&lt;th&gt;email&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;10.0&lt;/td&gt;
&lt;td&gt;20.0&lt;/td&gt;
&lt;td&gt;2026&lt;/td&gt;
&lt;td&gt;&lt;a href="mailto:ada@x.io"&gt;ada@x.io&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Projection gem&lt;/strong&gt;&lt;/strong&gt; — a Reformat compiles to one &lt;code&gt;select&lt;/code&gt;, so N derived columns cost one plan node, not N chained &lt;code&gt;withColumn&lt;/code&gt; calls that stack projections.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Column expressions&lt;/strong&gt;&lt;/strong&gt; — writing &lt;code&gt;qty * price&lt;/code&gt; in the gem and letting Prophecy emit &lt;code&gt;F.col("qty") * F.col("price")&lt;/code&gt; keeps the visual and the code semantically identical.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Explicit aliases&lt;/strong&gt;&lt;/strong&gt; — every derived column is named, so the downstream schema is deterministic and design-time validation can catch typos.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Optimizer-friendly&lt;/strong&gt;&lt;/strong&gt; — a single flat projection lets Catalyst fuse expressions, which is why "one gem, one select" is both cleaner and faster than a pile of &lt;code&gt;withColumn&lt;/code&gt;s.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the transform is O(rows) with no shuffle; a projection is a narrow transformation, so it adds no stage boundary.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Spark SQL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — spark-sql&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Spark SQL transform and projection problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/spark-sql" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;ETL&lt;/span&gt;
&lt;span&gt;Topic — etl&lt;/span&gt;
&lt;strong&gt;ETL pipeline-design problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. The visual ↔ code round-trip on Git
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Every visual edit regenerates readable code, and the whole project lives in your Git repo
&lt;/h3&gt;

&lt;p&gt;The feature that sells Prophecy to a platform team is the &lt;strong&gt;round-trip&lt;/strong&gt;: the canvas and the code are two views of the same artifact, kept in sync bidirectionally, and both are stored in ordinary version control. Edit a gem and the code regenerates; edit the code and the canvas re-parses. This is the difference between a low-code tool you cannot audit and one that behaves like a normal software repository.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How the round-trip works.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Visual → code.&lt;/strong&gt; Every change on the canvas — adding a gem, editing an expression, rewiring an edge — immediately regenerates the corresponding function in the project's source files. There is no separate "export" step; the code &lt;em&gt;is&lt;/em&gt; the save format.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code → visual.&lt;/strong&gt; Because Prophecy parses the generated code back into the graph, an engineer can open &lt;code&gt;pipeline.py&lt;/code&gt; in their IDE, edit it, push, and see the canvas reflect the change. The code and diagram cannot drift apart.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Standard, readable output.&lt;/strong&gt; The generated PySpark/SQL is not obfuscated machine output; it reads like code a person wrote, with named functions per gem, so it survives a human code review.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Git is the backbone, not a bolt-on.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Projects are repositories.&lt;/strong&gt; A Prophecy project maps to a Git repo (GitHub, GitLab, Bitbucket, Azure DevOps, or Prophecy-managed Git). Pipelines, subgraphs, tests, and configs are all files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Branch, commit, PR, merge.&lt;/strong&gt; You develop on a branch, commit visual changes as normal diffs, open a pull request, get a review, and merge — the full software lifecycle, driven from a visual tool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CI/CD like any codebase.&lt;/strong&gt; Because the artifact is code, you plug the repo into your existing CI to run tests and into your existing CD to deploy, with no Prophecy-specific pipeline server in the critical path.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why "no lock-in" is a technical claim, not a slogan.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The generated Spark job runs on any Spark cluster with &lt;code&gt;spark-submit&lt;/code&gt;; the generated SQL runs through dbt. Neither needs Prophecy at execution time.&lt;/li&gt;
&lt;li&gt;If you stopped using Prophecy, you keep a working, readable, tested Spark/SQL codebase — the exit cost is "you lose the visual editor," not "your pipelines stop running."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5kegpdk6zpl62ozh4ks.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5kegpdk6zpl62ozh4ks.jpeg" alt="Iconographic Prophecy round-trip diagram — a visual canvas and a code file kept in sync by a two-way arrow, both committed to a Git branch with commit, PR and merge glyphs, and a note that the generated code runs without Prophecy." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — an Aggregate gem and its Git diff
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The clearest way to see the round-trip is to add one Aggregate gem and read the diff it produces. The gem groups orders by customer and sums the total; the generated function is a plain &lt;code&gt;groupBy(...).agg(...)&lt;/code&gt;, and the Git diff a reviewer sees is a single new named function — reviewable in seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Add an Aggregate gem that computes &lt;code&gt;revenue = sum(total)&lt;/code&gt; and &lt;code&gt;orders = count(*)&lt;/code&gt; per &lt;code&gt;customer&lt;/code&gt;, and show the generated function that lands in the commit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;20.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;5.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;linus&lt;/td&gt;
&lt;td&gt;12.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;functions&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;AggregateByCustomer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;groupBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
           &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;agg&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
               &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;total&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;alias&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;revenue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
               &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;alias&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;orders&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
           &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The Aggregate gem's group-by key (&lt;code&gt;customer&lt;/code&gt;) becomes &lt;code&gt;groupBy("customer")&lt;/code&gt;, and each aggregate expression becomes one &lt;code&gt;.agg(...)&lt;/code&gt; entry with an alias. Prophecy writes this as a single new function; the commit diff is &lt;code&gt;+ def AggregateByCustomer(...)&lt;/code&gt; plus one wiring line in the driver. A reviewer reads the diff exactly as they would review a hand-written Spark change — because it is one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;revenue&lt;/th&gt;
&lt;th&gt;orders&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;25.0&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;linus&lt;/td&gt;
&lt;td&gt;12.0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If a visual change does not produce a clean, reviewable code diff, you are drawing something the code cannot express cleanly — split it into more gems until each diff is one idea.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prophecy interview question on aggregation transforms
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; An interviewer asks you to compute, per &lt;code&gt;customer&lt;/code&gt;, the total revenue and the number of &lt;em&gt;distinct&lt;/em&gt; order days, then keep only customers with revenue over 100. Write the Spark aggregation this gem would generate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using groupBy with a filtered aggregate
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;functions&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;CustomerRollup&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;groupBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
           &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;agg&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
               &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;total&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;alias&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;revenue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
               &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;countDistinct&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;order_ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;alias&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;active_days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
           &lt;span class="p"&gt;)&lt;/span&gt;
           &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;revenue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;rows in&lt;/th&gt;
&lt;th&gt;sum(total)&lt;/th&gt;
&lt;th&gt;distinct days&lt;/th&gt;
&lt;th&gt;kept?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;140.0&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;linus&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;60.0&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;grace&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;220.0&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;groupBy("customer")&lt;/code&gt; collapses each customer's rows into one group, triggering one shuffle keyed on &lt;code&gt;customer&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;sum("total")&lt;/code&gt; accumulates revenue and &lt;code&gt;countDistinct(to_date(order_ts))&lt;/code&gt; counts unique calendar days, both computed in the same aggregation pass.&lt;/li&gt;
&lt;li&gt;The post-aggregation &lt;code&gt;filter(revenue &amp;gt; 100)&lt;/code&gt; is a &lt;strong&gt;HAVING&lt;/strong&gt; clause: it runs on the grouped rows, dropping &lt;code&gt;linus&lt;/code&gt; while keeping &lt;code&gt;ada&lt;/code&gt; and &lt;code&gt;grace&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Ordering matters — filtering after the aggregate means the predicate sees &lt;code&gt;revenue&lt;/code&gt;, a column that only exists post-group; a Filter gem placed before the Aggregate could not reference it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;revenue&lt;/th&gt;
&lt;th&gt;active_days&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;140.0&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;grace&lt;/td&gt;
&lt;td&gt;220.0&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Group-by shuffle&lt;/strong&gt;&lt;/strong&gt; — the aggregation repartitions rows by &lt;code&gt;customer&lt;/code&gt; so all of one customer's rows meet on one executor; this is the single shuffle the whole transform costs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Multiple aggregates, one pass&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;sum&lt;/code&gt; and &lt;code&gt;countDistinct&lt;/code&gt; are computed together, so two metrics cost one scan and one shuffle, not two.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;HAVING via post-filter&lt;/strong&gt;&lt;/strong&gt; — filtering after &lt;code&gt;agg&lt;/code&gt; is the Spark equivalent of SQL &lt;code&gt;HAVING&lt;/code&gt;; the predicate references the aggregated &lt;code&gt;revenue&lt;/code&gt;, which is only defined after the group.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;countDistinct semantics&lt;/strong&gt;&lt;/strong&gt; — counting distinct &lt;code&gt;to_date(order_ts)&lt;/code&gt; answers "how many active days," a different question from &lt;code&gt;count(*)&lt;/code&gt;; picking the right counter is the correctness crux.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — O(rows) to scan plus one O(rows) shuffle keyed on &lt;code&gt;customer&lt;/code&gt;; the final filter is O(groups), typically tiny.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;ETL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — etl&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;ETL aggregation and rollup problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Transform&lt;/span&gt;
&lt;span&gt;Topic — data-transformation&lt;/span&gt;
&lt;strong&gt;Group-by and HAVING transform problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-transformation" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Running on Spark &amp;amp; Databricks with Fabrics
&lt;/h2&gt;
&lt;h3&gt;
  
  
  A Fabric binds a pipeline to where it runs — Spark on a cluster, or SQL on a warehouse
&lt;/h3&gt;

&lt;p&gt;Authoring produces code; something has to run it. In Prophecy that something is a &lt;strong&gt;Fabric&lt;/strong&gt; — the named execution environment that holds the connection, credentials, and cluster or warehouse settings your pipeline runs against. The same canvas can point at different Fabrics (dev, staging, prod), and the project language decides what the generated code targets: &lt;strong&gt;Spark projects&lt;/strong&gt; emit PySpark/Scala that runs on a cluster, &lt;strong&gt;SQL projects&lt;/strong&gt; emit dbt models that run on a warehouse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What a Fabric is.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The execution environment.&lt;/strong&gt; A Fabric bundles &lt;em&gt;where&lt;/em&gt; and &lt;em&gt;how&lt;/em&gt; a pipeline runs: a Databricks workspace and cluster, a Spark-on-EMR/Dataproc/Livy endpoint, or a SQL warehouse (Databricks SQL, Snowflake, BigQuery) for SQL projects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Connection + credentials + compute.&lt;/strong&gt; It carries the workspace URL, auth token/secret, and cluster size, so switching from a small dev cluster to a large prod cluster is a Fabric swap, not a code change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interactive and scheduled.&lt;/strong&gt; During development the Fabric backs the live per-gem preview; for production you deploy the pipeline as a scheduled &lt;strong&gt;Databricks Job&lt;/strong&gt; or an &lt;strong&gt;Airflow&lt;/strong&gt; DAG that calls the same generated code.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Spark projects vs SQL projects.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Spark project → PySpark/Scala.&lt;/strong&gt; Gems compile to &lt;code&gt;DataFrame&lt;/code&gt; transformations; the job is submitted to a Spark cluster. Choose this for large-scale distributed transforms, complex joins, and ML feature pipelines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SQL project → dbt.&lt;/strong&gt; Gems compile to SQL models with &lt;code&gt;ref()&lt;/code&gt;/&lt;code&gt;source()&lt;/code&gt; dependencies that run in-warehouse via dbt. Choose this when the data already lives in the warehouse and you want warehouse-native, dbt-tested models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same authoring, different target.&lt;/strong&gt; The gem canvas looks the same either way; the project language decides the dialect and the engine, which is how one tool serves both Spark and SQL teams.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers probe about execution.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Development ≠ deployment.&lt;/strong&gt; Interactive runs use an attached cluster for fast feedback; production runs are scheduled jobs on right-sized compute. Conflating the two is a classic junior mistake.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cluster sizing is a Fabric concern.&lt;/strong&gt; Spilling, out-of-memory, and shuffle-partition tuning are properties of the Fabric's compute, not the visual pipeline — the generated Spark obeys the same &lt;code&gt;spark.sql.shuffle.partitions&lt;/code&gt; you would tune by hand.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F31407mzvsvk9nclfjeau.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F31407mzvsvk9nclfjeau.jpeg" alt="Iconographic Prophecy execution diagram — a Fabric config card connecting to a Databricks Spark cluster and a SQL warehouse, with a Spark project compiling to PySpark and a SQL project compiling to dbt models, deployed as scheduled jobs." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — the same join as Spark and as SQL
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; Because the project language chooses the dialect, the same visual Join gem generates PySpark in a Spark project and SQL in a SQL project. Seeing both makes the "same canvas, different engine" point concrete: an &lt;code&gt;orders&lt;/code&gt;-to-&lt;code&gt;customers&lt;/code&gt; inner join is one gem, two outputs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Join &lt;code&gt;orders&lt;/code&gt; to &lt;code&gt;customers&lt;/code&gt; on &lt;code&gt;customer_id&lt;/code&gt; to attach &lt;code&gt;region&lt;/code&gt;. Show what the Join gem generates in a Spark project versus a SQL project.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;customer_id&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;## Spark project (PySpark) — generated Join gem
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;JoinCustomers&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;on&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;how&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inner&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; \
                 &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;order_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- SQL project (dbt model) — generated Join gem&lt;/span&gt;
&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;order_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;region&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'orders'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;
&lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'customers'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
  &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; In the Spark project the Join gem emits &lt;code&gt;orders.join(customers, on="customer_id", how="inner")&lt;/code&gt; and a projection selecting &lt;code&gt;region&lt;/code&gt;. In the SQL project the identical gem emits an ANSI &lt;code&gt;JOIN&lt;/code&gt; with dbt &lt;code&gt;ref()&lt;/code&gt; calls that wire model dependencies. The canvas, the join key, and the join type are the same; only the target dialect and engine differ, decided by the project language and the Fabric.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;customer_id&lt;/th&gt;
&lt;th&gt;region&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Pick the project language by where the data lives and how big it is — Spark for distributed scale, SQL/dbt for warehouse-native models — then let the same gem canvas target either.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prophecy interview question on deduplicating with Spark SQL
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Orders arrive with duplicate rows per &lt;code&gt;order_id&lt;/code&gt; (retries), each with an &lt;code&gt;updated_at&lt;/code&gt;. Keep exactly one row per &lt;code&gt;order_id&lt;/code&gt; — the latest &lt;code&gt;updated_at&lt;/code&gt; — using Spark SQL, the way a Deduplicate gem would generate it. How do you make it deterministic?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using ROW_NUMBER over a partition
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;ranked&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="k"&gt;select&lt;/span&gt;
    &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;row_number&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="n"&gt;over&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="k"&gt;partition&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="n"&gt;order_id&lt;/span&gt;
      &lt;span class="k"&gt;order&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="n"&gt;updated_at&lt;/span&gt; &lt;span class="k"&gt;desc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ingest_seq&lt;/span&gt; &lt;span class="k"&gt;desc&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;rn&lt;/span&gt;
  &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rn&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;ranked&lt;/span&gt;
&lt;span class="k"&gt;where&lt;/span&gt; &lt;span class="n"&gt;rn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;updated_at&lt;/th&gt;
&lt;th&gt;ingest_seq&lt;/th&gt;
&lt;th&gt;rn&lt;/th&gt;
&lt;th&gt;kept&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;12:30&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;12:30&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;11:00&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;10:00&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;partition by order_id&lt;/code&gt; groups the duplicate rows for each order together so ranking is per-key.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;order by updated_at desc&lt;/code&gt; puts the freshest row first; the tie-break &lt;code&gt;ingest_seq desc&lt;/code&gt; makes the winner &lt;strong&gt;deterministic&lt;/strong&gt; when two rows share the same &lt;code&gt;updated_at&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;row_number()&lt;/code&gt; assigns 1 to the winner of each partition, 2..N to the rest.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;where rn = 1&lt;/code&gt; keeps exactly one row per &lt;code&gt;order_id&lt;/code&gt;, and &lt;code&gt;except (rn)&lt;/code&gt; drops the helper column so the output schema matches the input.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;updated_at&lt;/th&gt;
&lt;th&gt;ingest_seq&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;12:30&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;11:00&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Window partition&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;partition by order_id&lt;/code&gt; turns "dedupe per key" into a ranking problem solved inside each partition, no self-join required.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Deterministic tie-break&lt;/strong&gt;&lt;/strong&gt; — adding &lt;code&gt;ingest_seq desc&lt;/code&gt; to the &lt;code&gt;order by&lt;/code&gt; removes the nondeterminism of ties, so re-running yields the same survivor every time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;row_number vs rank&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;row_number&lt;/code&gt; guarantees a single winner (no ties at rank 1), which is exactly the "keep one" requirement; &lt;code&gt;rank&lt;/code&gt; could keep several.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;except (rn)&lt;/strong&gt;&lt;/strong&gt; — projecting away the helper column keeps the deduplicated output schema-identical to the source, so downstream gems are unaffected.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — one shuffle to partition by &lt;code&gt;order_id&lt;/code&gt; plus an in-partition sort, O(rows log rows) per partition; a Deduplicate gem generates exactly this plan.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Spark SQL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — spark-sql&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Window-function and dedup problems in Spark SQL&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/spark-sql" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Optimization&lt;/span&gt;
&lt;span&gt;Topic — optimization&lt;/span&gt;
&lt;strong&gt;Shuffle and partition tuning problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/optimization" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Reusable subgraphs, tests &amp;amp; column-level lineage
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Package gems into reusable subgraphs, unit-test them, and trace every column
&lt;/h3&gt;

&lt;p&gt;The three features that make Prophecy an &lt;em&gt;enterprise&lt;/em&gt; tool rather than a personal one are &lt;strong&gt;reuse, testing, and lineage&lt;/strong&gt;. A reusable subgraph turns a proven set of gems into a shareable component; gem-level unit tests pin behaviour in CI; and column-level lineage answers "if I change this field, what breaks?" across every pipeline in the org. These are the capabilities a legacy visual ETL suite either lacks or hides behind an opaque metadata store.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reusable subgraphs.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A subgraph is a group of gems as one node.&lt;/strong&gt; You select several gems — say, the five that standardize a customer record — and package them into a &lt;strong&gt;reusable subgraph&lt;/strong&gt; with defined input and output ports. It appears on the canvas as a single box.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build once, reuse everywhere.&lt;/strong&gt; Every pipeline that needs "standardize customer" drops in the same subgraph; a fix to the subgraph propagates to all consumers, so business logic is DRY instead of copy-pasted across 40 pipelines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parameterizable.&lt;/strong&gt; Subgraphs (and custom gems built with the Gem Builder) take configuration, so one component adapts to slightly different inputs without forking.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Tests that run in CI.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gem-level unit tests.&lt;/strong&gt; You attach sample input rows and expected output rows to a gem; Prophecy generates a real unit test (pytest for PySpark, dbt tests for SQL) that runs in your CI on every pull request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail the build, not production.&lt;/strong&gt; Because tests are code in the repo, a transform regression fails the PR check — the same guardrail a hand-written Spark codebase has, now automatic from the visual definition.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data-quality expectations.&lt;/strong&gt; Beyond unit tests, expectation checks (not-null, uniqueness, accepted ranges) can gate a load so bad data fails loudly instead of corrupting a table.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Column-level lineage.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Field-to-field tracing.&lt;/strong&gt; Prophecy captures how each output column was derived from input columns across gems and pipelines, so you can trace &lt;code&gt;revenue&lt;/code&gt; back through every transform to its source columns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Impact analysis.&lt;/strong&gt; Before renaming or dropping a source column, lineage shows every downstream pipeline and report that depends on it — the difference between a safe change and a Monday-morning outage.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwc6m23udee8zinv502rn.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwc6m23udee8zinv502rn.jpeg" alt="Iconographic Prophecy reuse diagram — a group of gems packaged as a reusable subgraph component, a gem unit test with sample input and expected output, and a column-level lineage graph tracing a field across pipelines." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — a reusable "standardize customer" subgraph
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday reuse pattern is a subgraph that encapsulates a few cleaning gems behind one input and one output port. Its generated code is a function that takes a &lt;code&gt;DataFrame&lt;/code&gt; and returns a &lt;code&gt;DataFrame&lt;/code&gt;, so any pipeline can call it — the visual box and the reusable function are, again, the same thing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Package "trim + lowercase email, uppercase name, drop test rows" into a reusable subgraph and show the generated function that pipelines call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; A raw &lt;code&gt;customers&lt;/code&gt; DataFrame with mixed-case and test rows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;functions&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;StandardizeCustomer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))))&lt;/span&gt;
           &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;upper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;
           &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;~&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;endswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;@test.local&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The subgraph's three gems become three chained transformations inside one function: normalize &lt;code&gt;email&lt;/code&gt;, uppercase &lt;code&gt;name&lt;/code&gt;, and drop internal test accounts. Because it is exposed as a single &lt;code&gt;DataFrame → DataFrame&lt;/code&gt; function with one input port, every pipeline reuses it by wiring one box — and a fix here updates all of them at once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;name&lt;/th&gt;
&lt;th&gt;email&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ADA&lt;/td&gt;
&lt;td&gt;&lt;a href="mailto:ada@x.io"&gt;ada@x.io&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GRACE&lt;/td&gt;
&lt;td&gt;&lt;a href="mailto:grace@y.io"&gt;grace@y.io&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If the same three-or-more gems appear in two pipelines, promote them to a reusable subgraph — you trade a one-time packaging step for org-wide consistency and a single place to fix bugs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prophecy interview question on testing a transform
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You must guarantee that the "standardize customer" transform always lowercases email and never lets a &lt;code&gt;@test.local&lt;/code&gt; row through. Write the unit test a Prophecy gem test generates so this is enforced in CI, and explain what a failure would catch.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using a gem unit test with expected output
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;mypipeline.gems&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;StandardizeCustomer&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_standardize_customer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;src&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createDataFrame&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;[(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Ada&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; ADA@X.IO &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bot&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bot@test.local&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)],&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;StandardizeCustomer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;collect&lt;/span&gt;&lt;span class="p"&gt;()}&lt;/span&gt;

    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ADA&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ada@x.io&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;          &lt;span class="c1"&gt;# Bot dropped, Ada normalized
&lt;/span&gt;    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;                       &lt;span class="c1"&gt;# test row filtered out
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;input (name, email)&lt;/th&gt;
&lt;th&gt;after normalize&lt;/th&gt;
&lt;th&gt;after filter&lt;/th&gt;
&lt;th&gt;in expected?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ada, " &lt;a href="mailto:ADA@X.IO"&gt;ADA@X.IO&lt;/a&gt; "&lt;/td&gt;
&lt;td&gt;ADA, &lt;a href="mailto:ada@x.io"&gt;ada@x.io&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;kept&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bot, &lt;a href="mailto:bot@test.local"&gt;bot@test.local&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;BOT, &lt;a href="mailto:bot@test.local"&gt;bot@test.local&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;dropped&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The test builds a tiny &lt;code&gt;DataFrame&lt;/code&gt; with one good row and one &lt;code&gt;@test.local&lt;/code&gt; row, then calls the exact gem function under test.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;StandardizeCustomer&lt;/code&gt; trims and lowercases the email, uppercases the name, and filters the &lt;code&gt;@test.local&lt;/code&gt; row out.&lt;/li&gt;
&lt;li&gt;The assertion pins the &lt;em&gt;whole&lt;/em&gt; expected output — one row, &lt;code&gt;{"ADA": "ada@x.io"}&lt;/code&gt; — so any regression (a missing &lt;code&gt;lower&lt;/code&gt;, a broken filter) fails the assertion.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;out.count() == 1&lt;/code&gt; independently guards the filter, so a change that lets test rows through fails even if the surviving row still looks right.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;check&lt;/th&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;normalized email&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ada@x.io&lt;/code&gt; ✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;test row filtered&lt;/td&gt;
&lt;td&gt;count == 1 ✓&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Gem-level unit test&lt;/strong&gt;&lt;/strong&gt; — testing one gem function in isolation makes failures point at exactly one transform, not a whole pipeline run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Expected-output assertion&lt;/strong&gt;&lt;/strong&gt; — pinning the full output (not just "no error") is what turns a smoke test into a real regression guard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;CI enforcement&lt;/strong&gt;&lt;/strong&gt; — because the test is code in the repo, it runs on every pull request, so a visual edit that breaks the contract fails the build, not production.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Independent invariants&lt;/strong&gt;&lt;/strong&gt; — asserting both the value and the row count catches two distinct failure modes (bad normalization vs a broken filter) separately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — a two-row test runs in milliseconds locally, so it is cheap enough to attach to every gem that carries business rules.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Transform&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — data-transformation&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Reusable transform and cleaning problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-transformation" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Quality&lt;/span&gt;
&lt;span&gt;Topic — data-quality&lt;/span&gt;
&lt;strong&gt;Unit-test and data-validation problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-quality" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — Prophecy recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Minimal Spark pipeline (Source → Reformat → Target).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;pipeline&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;src&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;table&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sales.orders&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;total&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;qty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;overwrite&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;saveAsTable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sales.orders_enriched&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Generated gem function shape (one function per gem).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;Reformat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;qty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;alias&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;total&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Aggregate gem → groupBy/agg.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;Aggregate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;groupBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;agg&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;total&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;alias&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;revenue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;SQL project (dbt model) gem.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;order_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;region&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'orders'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;
&lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="p"&gt;{{&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'customers'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Reusable subgraph signature.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;StandardizeCustomer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;   &lt;span class="c1"&gt;# one input port, one output port
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;in0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Gem unit test skeleton.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_gem&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;src&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createDataFrame&lt;/span&gt;&lt;span class="p"&gt;([(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;10.0&lt;/span&gt;&lt;span class="p"&gt;)],&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;qty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Reformat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;collect&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;total&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mf"&gt;20.0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Choosing project language.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Project type&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Large distributed transforms, complex joins, ML features&lt;/td&gt;
&lt;td&gt;Spark (PySpark/Scala)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data already in the warehouse, want dbt-tested models&lt;/td&gt;
&lt;td&gt;SQL (dbt)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Migrating Informatica/Alteryx mappings to open code&lt;/td&gt;
&lt;td&gt;Spark or SQL by target engine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analysts contributing reviewable pipelines&lt;/td&gt;
&lt;td&gt;either — both emit Git-tracked code&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is Prophecy?
&lt;/h3&gt;

&lt;p&gt;Prophecy is a low-code data engineering platform: a visual, drag-and-drop canvas for building data pipelines that &lt;strong&gt;compiles to open-source code&lt;/strong&gt;. You wire together &lt;strong&gt;gems&lt;/strong&gt; (transform nodes) on a graph, and Prophecy generates readable PySpark or Scala Spark for Spark pipelines, or dbt-style SQL models for warehouse pipelines, and commits that code to your own Git repository. It is a design-and-generate layer, not a runtime you send data through.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Prophecy lock me into its runtime?
&lt;/h3&gt;

&lt;p&gt;No — that is the central design choice. The artifact Prophecy produces is standard Spark or SQL code in your Git repo; the Spark runs on any Spark cluster (Databricks, EMR, Dataproc) via &lt;code&gt;spark-submit&lt;/code&gt;, and the SQL runs through dbt on your warehouse. Neither needs Prophecy at execution time, so if you stopped using the visual editor you would still have a working, readable, tested codebase.&lt;/p&gt;

&lt;h3&gt;
  
  
  What are gems in Prophecy?
&lt;/h3&gt;

&lt;p&gt;Gems are the building blocks of a pipeline — each gem is one transform node on the canvas. Built-in gems cover the standard operations (Source, Target, Reformat, Aggregate, Join, Filter, OrderBy, Deduplicate, SetOperation, FlattenSchema), and each gem compiles to exactly one function in the generated code. Platform teams can also build &lt;strong&gt;custom gems&lt;/strong&gt; with the Gem Builder to encode company-specific logic as reusable drag-and-drop boxes.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does the visual-to-code round-trip work?
&lt;/h3&gt;

&lt;p&gt;Every edit on the visual canvas immediately regenerates the corresponding code, and because Prophecy also parses code back into the graph, an engineer can edit the generated file in an IDE and see the canvas update. The two views cannot drift apart. Since the code is the save format and lives in Git, you get branches, commits, pull requests, reviews, and CI/CD exactly like any software repository.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Prophecy support SQL as well as Spark?
&lt;/h3&gt;

&lt;p&gt;Yes. A &lt;strong&gt;Spark project&lt;/strong&gt; compiles gems to PySpark/Scala &lt;code&gt;DataFrame&lt;/code&gt; transformations that run on a Spark cluster, while a &lt;strong&gt;SQL project&lt;/strong&gt; compiles the same style of gem canvas to dbt SQL models that run in-warehouse (Databricks SQL, Snowflake, BigQuery). The authoring experience is the same; the project language and the &lt;strong&gt;Fabric&lt;/strong&gt; (execution environment) decide the dialect and engine.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can Prophecy migrate legacy Informatica or Alteryx pipelines?
&lt;/h3&gt;

&lt;p&gt;Yes — a common enterprise use case is converting legacy visual ETL (Informatica, Alteryx, Ab Initio, SSIS, DataStage) into open Spark or SQL. Prophecy's migration tooling maps the legacy mappings/workflows onto gems and generates the equivalent PySpark or dbt code in Git, so you exit the proprietary runtime while keeping a visual editor and a modern, testable, version-controlled codebase.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every Prophecy idea above, from the projection Reformat gem to the group-by rollup, the ROW_NUMBER dedup, and the unit-tested reusable subgraph, maps to a hands-on practice room where you write the Spark or SQL against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "what code would that gem generate?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/spark-sql" rel="noopener noreferrer"&gt;Practice Spark SQL problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;ETL pipeline drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Bruin: SQL + Python Pipelines in One Framework With Built-In Quality Checks</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Sun, 27 Sep 2026 17:43:16 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/bruin-sql-python-pipelines-in-one-framework-with-built-in-quality-checks-7ip</link>
      <guid>https://dev.to/gowthampotureddi/bruin-sql-python-pipelines-in-one-framework-with-built-in-quality-checks-7ip</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;bruin&lt;/code&gt;&lt;/strong&gt; is the open-source data pipeline framework that collapses four tools into one: it runs your SQL transforms, executes your Python jobs, ingests data from external sources, and enforces data-quality checks — all from a single Go binary and a folder of files you keep in version control. You do not wire dbt to Airbyte to Great Expectations to an orchestrator and hope the seams hold. You write &lt;em&gt;assets&lt;/em&gt; — SQL files, Python files, or small YAML files — each carrying its own metadata in a comment block, and you type &lt;code&gt;bruin run&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That is a different shape from the modern stack most teams assembled over the last five years, where transformation, ingestion, testing, and scheduling are four separate products glued together with CI scripts. This guide walks through the four ideas an interviewer will actually probe — the asset model with YAML-in-comment metadata, materialization strategies, built-in and custom quality checks, and ingestion plus the dependency-driven DAG — and pairs each with a Solution-Tail interview answer: code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm6goz7ds8rdlogxlq2s2.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm6goz7ds8rdlogxlq2s2.jpeg" alt="PipeCode blog header for Bruin — bold white headline 'Bruin: SQL + Python' with subtitle 'one framework · built-in quality checks' and a stylised asset-DAG-to-warehouse scene on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;pipeline-design practice library →&lt;/a&gt;, rehearse the load-shape decisions on the &lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;ETL practice set →&lt;/a&gt;, and harden your test suite on the &lt;a href="https://pipecode.ai/explore/practice/topic/data-quality" rel="noopener noreferrer"&gt;data-quality practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why Bruin puts the whole pipeline in one framework&lt;/li&gt;
&lt;li&gt;Assets: SQL, Python &amp;amp; YAML-in-comment metadata&lt;/li&gt;
&lt;li&gt;Materialization: how a SELECT becomes a table&lt;/li&gt;
&lt;li&gt;Built-in &amp;amp; custom quality checks&lt;/li&gt;
&lt;li&gt;Ingestion with ingestr &amp;amp; multi-destination DAGs&lt;/li&gt;
&lt;li&gt;Cheat sheet — Bruin recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why Bruin puts the whole pipeline in one framework
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Bruin is one tool for SQL, Python, ingestion, and quality — that single fact decides where it fits
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;Bruin treats every step of a pipeline — extract, transform, and test — as an &lt;em&gt;asset&lt;/em&gt; in one project, so you stop stitching four products together and run the whole graph with &lt;code&gt;bruin run&lt;/code&gt;&lt;/strong&gt;. Everything that makes Bruin attractive to a data engineering team follows from that. There is no separate ingestion service, no separate test framework, no separate scheduler required to make a pipeline correct; a Bruin project is a folder of asset files plus a &lt;code&gt;pipeline.yml&lt;/code&gt; that the CLI reads, orders, executes, and validates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What lives inside one Bruin project.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ingestion.&lt;/strong&gt; An &lt;code&gt;ingestr&lt;/code&gt; asset pulls data from an external source (an API, a database, a bucket) into a destination, with no connector code to write.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transformation.&lt;/strong&gt; A SQL asset (&lt;code&gt;bq.sql&lt;/code&gt;, &lt;code&gt;sf.sql&lt;/code&gt;, &lt;code&gt;duckdb.sql&lt;/code&gt;, &lt;code&gt;pg.sql&lt;/code&gt;, and more) or a Python asset runs your business logic and materializes a table or view.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality.&lt;/strong&gt; Column-level and custom checks are declared &lt;em&gt;in the same file&lt;/em&gt; as the asset they guard, and run right after it produces data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Orchestration.&lt;/strong&gt; Bruin reads each asset's &lt;code&gt;depends&lt;/code&gt; list, builds a DAG, and runs assets in topological order — no external orchestrator needed for a single pipeline.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Where Bruin sits against the alternatives.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;vs dbt.&lt;/strong&gt; dbt is transform-only: it compiles and runs SQL (and, more recently, Python models) but leaves ingestion and orchestration to other tools. Bruin covers the same SQL-modelling ground &lt;em&gt;plus&lt;/em&gt; native Python assets, &lt;code&gt;ingestr&lt;/code&gt;-based ingestion, and a built-in runner. If your answer to "how does raw data arrive?" is "a different tool," Bruin folds that step in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vs dbt + Airbyte + Great Expectations + Airflow.&lt;/strong&gt; The classic modern stack is four products and the glue between them. Bruin puts ingestion, transformation, quality, and single-pipeline scheduling behind one CLI and one config language, so there are fewer moving parts to break.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vs a hand-rolled Python stack.&lt;/strong&gt; A pile of scripts has no schema for metadata, no dependency graph, no first-class checks, and no consistent run semantics. Bruin gives you all four without a platform to operate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you say &lt;strong&gt;"Bruin is SQL + Python + ingestion + quality in one framework"&lt;/strong&gt; rather than "another dbt"? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you place dbt as &lt;strong&gt;"transform-only"&lt;/strong&gt; and Bruin as &lt;strong&gt;"end-to-end assets"&lt;/strong&gt; unprompted? — required framing.&lt;/li&gt;
&lt;li&gt;Do you mention that &lt;strong&gt;quality checks live next to the asset and can block downstream&lt;/strong&gt; — not in a separate suite? — the whole point.&lt;/li&gt;
&lt;li&gt;Do you note that &lt;strong&gt;&lt;code&gt;bruin run&lt;/code&gt; needs no external orchestrator&lt;/strong&gt; for a single pipeline, while still deploying to Airflow when you want it? — senior signal.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — one asset that materializes a table
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The canonical Bruin "hello world" is a single SQL file whose top comment block is YAML metadata and whose body is a &lt;code&gt;SELECT&lt;/code&gt;. Bruin reads the metadata to learn the asset's name, platform, and how to materialize it, then wraps your query in the right DDL. You never write &lt;code&gt;CREATE TABLE&lt;/code&gt; — you declare &lt;code&gt;materialization: type: table&lt;/code&gt; and Bruin does it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Write one Bruin SQL asset that turns a two-row &lt;code&gt;SELECT&lt;/code&gt; into a DuckDB table &lt;code&gt;dashboard.hello&lt;/code&gt;, with no &lt;code&gt;CREATE TABLE&lt;/code&gt; written by you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;what you write&lt;/th&gt;
&lt;th&gt;what Bruin uses it for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;@bruin&lt;/code&gt; comment block&lt;/td&gt;
&lt;td&gt;asset metadata (name, type, materialization)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;SELECT&lt;/code&gt; body&lt;/td&gt;
&lt;td&gt;the query whose result becomes the table&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: dashboard.hello
type: duckdb.sql
materialization:
  type: table
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'ada'&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;
&lt;span class="k"&gt;union&lt;/span&gt; &lt;span class="k"&gt;all&lt;/span&gt;
&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'linus'&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The block between &lt;code&gt;/* @bruin&lt;/code&gt; and &lt;code&gt;@bruin */&lt;/code&gt; is parsed as YAML: &lt;code&gt;name: dashboard.hello&lt;/code&gt; becomes the target &lt;code&gt;schema.table&lt;/code&gt;, &lt;code&gt;type: duckdb.sql&lt;/code&gt; tells Bruin to run the body as DuckDB SQL, and &lt;code&gt;materialization: type: table&lt;/code&gt; tells it to persist the result. When you run &lt;code&gt;bruin run assets/hello.sql&lt;/code&gt;, Bruin generates a &lt;code&gt;create+replace&lt;/code&gt; statement (the default table strategy), executes your &lt;code&gt;SELECT&lt;/code&gt;, and stores the rows in &lt;code&gt;dashboard.hello&lt;/code&gt;. The SQL body stays a plain &lt;code&gt;SELECT&lt;/code&gt; — all the persistence logic lives in the metadata.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Bruin created&lt;/th&gt;
&lt;th&gt;value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;table&lt;/td&gt;
&lt;td&gt;&lt;code&gt;dashboard.hello&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;columns&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;id&lt;/code&gt;, &lt;code&gt;name&lt;/code&gt; (inferred from the &lt;code&gt;SELECT&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;rows&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;strategy applied&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;create+replace&lt;/code&gt; (table default)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If a step can be expressed as "metadata plus a body," it is a Bruin asset — a SQL query, a Python script, or an ingestion spec differ only in the body, never in the project shape.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Assets: SQL, Python &amp;amp; YAML-in-comment metadata
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The asset — code plus its YAML metadata in one file — is Bruin's entire mental model
&lt;/h3&gt;

&lt;p&gt;Bruin has exactly one core abstraction you compose, and an interviewer who asks "walk me through Bruin's model" wants it crisp: &lt;strong&gt;everything is an asset, and an asset is code with a block of YAML metadata attached to it&lt;/strong&gt;. Learn the three file shapes and where the metadata goes, and the whole framework snaps into focus.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The three asset file shapes.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SQL asset (&lt;code&gt;.sql&lt;/code&gt;).&lt;/strong&gt; The metadata lives between &lt;code&gt;/* @bruin&lt;/code&gt; and &lt;code&gt;@bruin */&lt;/code&gt; markers at the top of the file; the SQL query body follows underneath. Definition and query stay in one file — Bruin will &lt;em&gt;not&lt;/em&gt; let you split them across a sibling &lt;code&gt;.asset.yml&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Python asset (&lt;code&gt;.py&lt;/code&gt;).&lt;/strong&gt; The metadata lives between &lt;code&gt;"""@bruin&lt;/code&gt; and &lt;code&gt;@bruin"""&lt;/code&gt; docstring markers; the Python script follows. Same file, same idea.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;YAML asset (&lt;code&gt;&amp;lt;name&amp;gt;.asset.yml&lt;/code&gt;).&lt;/strong&gt; A standalone YAML file with &lt;em&gt;only&lt;/em&gt; metadata and no code body — used for asset types that have no inline code, such as &lt;code&gt;ingestr&lt;/code&gt;, &lt;code&gt;sensor&lt;/code&gt;, and &lt;code&gt;seed&lt;/code&gt;. The &lt;code&gt;.asset.yml&lt;/code&gt; suffix is required; a plain &lt;code&gt;.yml&lt;/code&gt; is treated as config and ignored.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The metadata keys that matter.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;name&lt;/code&gt;.&lt;/strong&gt; The asset's identity, following the &lt;code&gt;schema.table&lt;/code&gt; convention (dots separate segments). It is optional — if omitted, Bruin infers it from the file path relative to &lt;code&gt;assets/&lt;/code&gt;, so &lt;code&gt;assets/analytics/orders.sql&lt;/code&gt; becomes &lt;code&gt;analytics.orders&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;type&lt;/code&gt;.&lt;/strong&gt; How the asset executes: &lt;code&gt;bq.sql&lt;/code&gt;, &lt;code&gt;sf.sql&lt;/code&gt;, &lt;code&gt;pg.sql&lt;/code&gt;, &lt;code&gt;duckdb.sql&lt;/code&gt;, &lt;code&gt;python&lt;/code&gt;, &lt;code&gt;ingestr&lt;/code&gt;, &lt;code&gt;r&lt;/code&gt;, and more. The type binds the asset to a platform.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;depends&lt;/code&gt;.&lt;/strong&gt; The list of upstream assets this one waits for. This list &lt;em&gt;is&lt;/em&gt; the edge set of the DAG — Bruin runs an asset only after every asset in its &lt;code&gt;depends&lt;/code&gt; has succeeded.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;materialization&lt;/code&gt;, &lt;code&gt;columns&lt;/code&gt;, &lt;code&gt;custom_checks&lt;/code&gt;.&lt;/strong&gt; How the result is persisted, the typed columns with their quality checks, and any SQL-expressed custom checks — all covered in the next sections.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why the asset is the powerful unit.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Because metadata sits next to the code, everything about a table — its owner, its dependencies, its tests — is readable in one file and versioned in one commit.&lt;/li&gt;
&lt;li&gt;Because SQL and Python are both just assets, a single pipeline can mix a &lt;code&gt;duckdb.sql&lt;/code&gt; staging model and a &lt;code&gt;python&lt;/code&gt; enrichment job in the same DAG, with a dependency edge between them.&lt;/li&gt;
&lt;li&gt;Because &lt;code&gt;name&lt;/code&gt; follows &lt;code&gt;schema.table&lt;/code&gt;, the asset graph mirrors your warehouse layout, and Bruin's VS Code extension can render lineage straight from the &lt;code&gt;depends&lt;/code&gt; edges.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5gvpa1ts3bma4lgzoeqn.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5gvpa1ts3bma4lgzoeqn.jpeg" alt="Iconographic Bruin asset-model diagram — a SQL asset with a @bruin comment header, a Python asset with a @bruin docstring, and a standalone YAML asset, each carrying name / type / depends metadata that Bruin assembles into a DAG." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — a Python asset that depends on a SQL asset
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; Real pipelines mix languages. Here a SQL asset builds a &lt;code&gt;staging.players&lt;/code&gt; table, and a Python asset depends on it to compute something SQL is awkward at. The &lt;code&gt;depends&lt;/code&gt; edge is what makes Bruin run the SQL asset first and the Python asset second, in one command.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Define a &lt;code&gt;duckdb.sql&lt;/code&gt; asset &lt;code&gt;staging.players&lt;/code&gt; and a &lt;code&gt;python&lt;/code&gt; asset &lt;code&gt;analytics.player_features&lt;/code&gt; that depends on it, so &lt;code&gt;bruin run&lt;/code&gt; executes them in the right order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; Two files under &lt;code&gt;assets/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: staging.players
type: duckdb.sql
materialization:
  type: table
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="s1"&gt;'ada'&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;rating&lt;/span&gt;
&lt;span class="k"&gt;union&lt;/span&gt; &lt;span class="k"&gt;all&lt;/span&gt;
&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="s1"&gt;'linus'&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;51&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;rating&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;@bruin
name: analytics.player_features
type: python
depends:
  - staging.players
@bruin&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;computing features from staging.players&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;  &lt;span class="c1"&gt;# a real asset would read the upstream table and write a result
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The SQL file's &lt;code&gt;@bruin&lt;/code&gt; block names it &lt;code&gt;staging.players&lt;/code&gt; and materializes it as a table. The Python file's docstring block names it &lt;code&gt;analytics.player_features&lt;/code&gt;, sets &lt;code&gt;type: python&lt;/code&gt;, and lists &lt;code&gt;staging.players&lt;/code&gt; under &lt;code&gt;depends&lt;/code&gt;. When you run &lt;code&gt;bruin run&lt;/code&gt;, Bruin parses both files, sees the edge &lt;code&gt;staging.players → analytics.player_features&lt;/code&gt;, topologically sorts them, and executes the SQL asset before the Python asset. Neither file references the other's code — the &lt;code&gt;depends&lt;/code&gt; metadata alone encodes the order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;asset&lt;/th&gt;
&lt;th&gt;type&lt;/th&gt;
&lt;th&gt;runs when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;staging.players&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;duckdb.sql&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;first (no upstream)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;analytics.player_features&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;python&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;after &lt;code&gt;staging.players&lt;/code&gt; succeeds&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; One file = one asset. Put the language in &lt;code&gt;type&lt;/code&gt;, put the order in &lt;code&gt;depends&lt;/code&gt;, and let Bruin, not a scheduler you configure separately, decide execution order.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bruin interview question on the asset model
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; An interviewer gives you a folder of SQL and Python files and asks: how does Bruin know an asset's table name, which platform runs it, and what must complete before it — without any central manifest listing all of this? Show a SQL asset that answers all three, and explain where each answer comes from.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using inline YAML metadata and path-based name inference
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: analytics.daily_revenue
type: bq.sql
owner: data-team@acme.com
depends:
  - staging.orders
  - staging.refunds
materialization:
  type: table
  strategy: create+replace
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt;
  &lt;span class="n"&gt;order_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;gross&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;refund&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;refunds&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;refund&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;net&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;staging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;
&lt;span class="k"&gt;left&lt;/span&gt; &lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="n"&gt;staging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;refunds&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;order_date&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;group&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="n"&gt;order_date&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;question&lt;/th&gt;
&lt;th&gt;answer&lt;/th&gt;
&lt;th&gt;where it comes from&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;table name?&lt;/td&gt;
&lt;td&gt;&lt;code&gt;analytics.daily_revenue&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the &lt;code&gt;name&lt;/code&gt; key (or the file path if omitted)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;which platform?&lt;/td&gt;
&lt;td&gt;BigQuery&lt;/td&gt;
&lt;td&gt;&lt;code&gt;type: bq.sql&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;what runs first?&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;staging.orders&lt;/code&gt;, &lt;code&gt;staging.refunds&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;the &lt;code&gt;depends&lt;/code&gt; list&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Bruin scans every file under &lt;code&gt;assets/&lt;/code&gt; and parses each file's inline &lt;code&gt;@bruin&lt;/code&gt; block — there is &lt;strong&gt;no&lt;/strong&gt; central manifest; the metadata is distributed into the files themselves.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;name: analytics.daily_revenue&lt;/code&gt; sets the target &lt;code&gt;schema.table&lt;/code&gt;; had it been omitted, Bruin would have inferred it from the path (&lt;code&gt;assets/analytics/daily_revenue.sql&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;type: bq.sql&lt;/code&gt; binds the asset to the BigQuery connection so the body runs as BigQuery SQL.&lt;/li&gt;
&lt;li&gt;Each entry in &lt;code&gt;depends&lt;/code&gt; becomes an inbound edge, so Bruin runs &lt;code&gt;staging.orders&lt;/code&gt; and &lt;code&gt;staging.refunds&lt;/code&gt; before this asset and fails fast if either upstream fails.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;resolved property&lt;/th&gt;
&lt;th&gt;value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;target&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;analytics.daily_revenue&lt;/code&gt; (BigQuery)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;upstream edges&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;staging.orders&lt;/code&gt;, &lt;code&gt;staging.refunds&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;execution position&lt;/td&gt;
&lt;td&gt;after both upstreams succeed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Inline metadata&lt;/strong&gt;&lt;/strong&gt; — because the YAML lives in the same file as the code, an asset is self-describing; there is no separate registry to keep in sync with the files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Name inference&lt;/strong&gt;&lt;/strong&gt; — the &lt;code&gt;schema.table&lt;/code&gt; convention (explicit or path-derived) means the asset graph mirrors the warehouse layout without redundant configuration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Type binds platform&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;type&lt;/code&gt; alone routes the body to BigQuery, Snowflake, DuckDB, or Python, so one project can span engines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;depends is the DAG&lt;/strong&gt;&lt;/strong&gt; — the dependency list is the single source of execution order; Bruin never guesses, and lineage is exact.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — parsing is O(assets) at startup; the DAG build is O(V + E) topological sort, negligible next to query runtime.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Pipelines&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — pipelines&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Pipeline-design and DAG-ordering problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;ETL&lt;/span&gt;
&lt;span&gt;Topic — etl&lt;/span&gt;
&lt;strong&gt;Extract-transform-load asset problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. Materialization: how a SELECT becomes a table
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Bruin wraps your SELECT in the load strategy you choose — from full refresh to merge and SCD2
&lt;/h3&gt;

&lt;p&gt;The feature that turns Bruin from "a SQL runner" into "a warehouse-loading framework" is &lt;strong&gt;materialization&lt;/strong&gt;: you write a plain &lt;code&gt;SELECT&lt;/code&gt;, and Bruin applies the DDL and DML needed to persist that result the way you asked. Say it in one breath: &lt;strong&gt;&lt;code&gt;type&lt;/code&gt; decides table-or-view, &lt;code&gt;strategy&lt;/code&gt; decides how rows hit the table&lt;/strong&gt;. Pick wrong and you get a data-quality bug; pick right and you never hand-write a &lt;code&gt;MERGE&lt;/code&gt; again.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The two materialization types.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;type: table&lt;/code&gt;.&lt;/strong&gt; Persist the query result as a table. This is the workhorse and the one that carries a &lt;code&gt;strategy&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;type: view&lt;/code&gt;.&lt;/strong&gt; Persist the query as a view — no data is copied, the query re-runs on read.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The strategies that matter (all on &lt;code&gt;type: table&lt;/code&gt;).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;create+replace&lt;/code&gt; (default).&lt;/strong&gt; Overwrite the whole table with the new result every run. A full refresh — simple and correct for small tables, expensive for large ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;delete+insert&lt;/code&gt;.&lt;/strong&gt; Incremental: needs an &lt;code&gt;incremental_key&lt;/code&gt;; Bruin loads the query into a temp table, deletes the target rows whose key values appear in the new batch, then inserts. Good for reprocessing a partition.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;truncate+insert&lt;/code&gt;.&lt;/strong&gt; Full replacement that keeps the existing table's schema, permissions, and indices — &lt;code&gt;TRUNCATE&lt;/code&gt; then &lt;code&gt;INSERT&lt;/code&gt;, no &lt;code&gt;DROP&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;append&lt;/code&gt;.&lt;/strong&gt; Only add the new rows, never overwrite — the correct choice for immutable event logs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;merge&lt;/code&gt;.&lt;/strong&gt; Upsert: mark columns with &lt;code&gt;primary_key: true&lt;/code&gt; and Bruin updates matched rows and inserts new ones. Use &lt;code&gt;update_on_merge&lt;/code&gt; to mark which columns get overwritten on a match, or &lt;code&gt;merge_sql&lt;/code&gt; for a custom expression like &lt;code&gt;GREATEST(target.col, source.col)&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;time_interval&lt;/code&gt;.&lt;/strong&gt; Incrementally load a time window; needs &lt;code&gt;incremental_key&lt;/code&gt; and &lt;code&gt;time_granularity&lt;/code&gt; (&lt;code&gt;date&lt;/code&gt; or &lt;code&gt;timestamp&lt;/code&gt;), and you pass &lt;code&gt;--start-date&lt;/code&gt; / &lt;code&gt;--end-date&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ddl&lt;/code&gt;.&lt;/strong&gt; Create an empty table from the &lt;code&gt;columns&lt;/code&gt; definition only — no query body — when you want the structure created exactly once.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;scd2_by_column&lt;/code&gt; / &lt;code&gt;scd2_by_time&lt;/code&gt;.&lt;/strong&gt; Maintain full Slowly-Changing-Dimension history with automatic &lt;code&gt;_valid_from&lt;/code&gt;, &lt;code&gt;_valid_until&lt;/code&gt;, and &lt;code&gt;_is_current&lt;/code&gt; columns, tracking changes by column diff or by a time-based incremental key.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgbz7ggh86qnzl0xa0fcx.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgbz7ggh86qnzl0xa0fcx.jpeg" alt="Iconographic Bruin materialization diagram — a SELECT query on the left, a materialization dial in the centre with create+replace, merge, and scd2 notches, and the resulting table on the right showing an upsert and an SCD2 history row." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — an incremental merge on a primary key
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The most common non-trivial materialization is a &lt;code&gt;merge&lt;/code&gt;: you have a dimension keyed by an id, rows change, and you want to update the changed ones and insert the new ones without duplicating. Bruin generates the &lt;code&gt;MERGE&lt;/code&gt; for you from the &lt;code&gt;columns&lt;/code&gt; metadata — you mark the key and the updatable columns and write a plain &lt;code&gt;SELECT&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Materialize a &lt;code&gt;dim.customers&lt;/code&gt; table that upserts on &lt;code&gt;customer_id&lt;/code&gt;, overwriting &lt;code&gt;email&lt;/code&gt; and &lt;code&gt;tier&lt;/code&gt; when a customer already exists.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer_id&lt;/th&gt;
&lt;th&gt;email&lt;/th&gt;
&lt;th&gt;tier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;ada@x&lt;/td&gt;
&lt;td&gt;gold&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;linus@x&lt;/td&gt;
&lt;td&gt;silver&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: dim.customers
type: bq.sql
materialization:
  type: table
  strategy: merge

columns:
  - name: customer_id
    type: integer
    primary_key: true
  - name: email
    type: string
    update_on_merge: true
  - name: tier
    type: string
    update_on_merge: true
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tier&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;staging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customers&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;strategy: merge&lt;/code&gt; tells Bruin to generate a &lt;code&gt;MERGE&lt;/code&gt; statement instead of overwriting the table. Bruin reads the &lt;code&gt;columns&lt;/code&gt; block: &lt;code&gt;customer_id&lt;/code&gt; is &lt;code&gt;primary_key: true&lt;/code&gt;, so it becomes the match key; &lt;code&gt;email&lt;/code&gt; and &lt;code&gt;tier&lt;/code&gt; carry &lt;code&gt;update_on_merge: true&lt;/code&gt;, so they are the columns overwritten with &lt;code&gt;source&lt;/code&gt; values when a row matches. On each run, Bruin runs your &lt;code&gt;SELECT&lt;/code&gt; as the source, matches on &lt;code&gt;customer_id&lt;/code&gt;, updates the matched rows' &lt;code&gt;email&lt;/code&gt;/&lt;code&gt;tier&lt;/code&gt;, and inserts the unmatched rows. Rows deleted at the source stay in the target — &lt;code&gt;merge&lt;/code&gt; never deletes, unlike &lt;code&gt;delete+insert&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer_id&lt;/th&gt;
&lt;th&gt;result on next run&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;5 (exists)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;email&lt;/code&gt;/&lt;code&gt;tier&lt;/code&gt; updated in place&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6 (exists)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;email&lt;/code&gt;/&lt;code&gt;tier&lt;/code&gt; updated in place&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9 (new)&lt;/td&gt;
&lt;td&gt;inserted&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Use &lt;code&gt;merge&lt;/code&gt; for mutable dimensions keyed by an id; use &lt;code&gt;append&lt;/code&gt; for immutable events; reach for &lt;code&gt;delete+insert&lt;/code&gt; only when the source truly removes rows and the target must mirror that.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bruin interview question on tracking history
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Product prices change over time and the business needs to answer "what was this product's price on any past date," not just its current price. Which Bruin materialization strategy gives you that with no hand-written history SQL, and what columns does it add?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using the scd2_by_column strategy
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: dim.product_catalog
type: bq.sql
materialization:
  type: table
  strategy: scd2_by_column

columns:
  - name: id
    type: integer
    primary_key: true
  - name: name
    type: string
  - name: price
    type: float
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;staging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;products&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;run&lt;/th&gt;
&lt;th&gt;id 1 price&lt;/th&gt;
&lt;th&gt;action taken&lt;/th&gt;
&lt;th&gt;resulting rows for id 1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 (full refresh)&lt;/td&gt;
&lt;td&gt;29.99&lt;/td&gt;
&lt;td&gt;insert current version&lt;/td&gt;
&lt;td&gt;1 current row&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;39.99&lt;/td&gt;
&lt;td&gt;expire old, insert new&lt;/td&gt;
&lt;td&gt;1 historical + 1 current&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;39.99 (unchanged)&lt;/td&gt;
&lt;td&gt;no change detected&lt;/td&gt;
&lt;td&gt;unchanged&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;strategy: scd2_by_column&lt;/code&gt; makes Bruin compare each incoming row against the current stored version, keyed by the &lt;code&gt;primary_key&lt;/code&gt; column &lt;code&gt;id&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;On the first run (&lt;code&gt;--full-refresh&lt;/code&gt;), each product is inserted as a current version with &lt;code&gt;_is_current = true&lt;/code&gt;, &lt;code&gt;_valid_from = now&lt;/code&gt;, and &lt;code&gt;_valid_until = 9999-12-31&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;When &lt;code&gt;price&lt;/code&gt; changes on run 2, Bruin marks the old row &lt;code&gt;_is_current = false&lt;/code&gt; and sets its &lt;code&gt;_valid_until&lt;/code&gt; to the change time, then inserts a new current row with the updated price.&lt;/li&gt;
&lt;li&gt;A row that disappears from the source is expired (&lt;code&gt;_is_current = false&lt;/code&gt;); an unchanged row is left alone. You never wrote a line of history-tracking SQL.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;id&lt;/th&gt;
&lt;th&gt;price&lt;/th&gt;
&lt;th&gt;_is_current&lt;/th&gt;
&lt;th&gt;_valid_from&lt;/th&gt;
&lt;th&gt;_valid_until&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;29.99&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;2024-01-01&lt;/td&gt;
&lt;td&gt;2024-01-02&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;39.99&lt;/td&gt;
&lt;td&gt;true&lt;/td&gt;
&lt;td&gt;2024-01-02&lt;/td&gt;
&lt;td&gt;9999-12-31&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;scd2_by_column&lt;/strong&gt;&lt;/strong&gt; — a declarative strategy that detects changes in any non-key column and versions the row, so history is a config choice, not bespoke SQL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Reserved validity columns&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;_valid_from&lt;/code&gt;, &lt;code&gt;_valid_until&lt;/code&gt;, and &lt;code&gt;_is_current&lt;/code&gt; are added and maintained by Bruin, giving you point-in-time queries with a simple &lt;code&gt;where _valid_from &amp;lt;= d and _valid_until &amp;gt; d&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;primary_key drives detection&lt;/strong&gt;&lt;/strong&gt; — the key identifies the entity across runs; changes to non-key columns trigger a new version, unchanged rows are skipped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Idempotent history&lt;/strong&gt;&lt;/strong&gt; — re-running with the same source inserts nothing new, because no change is detected — retries are safe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the change comparison is O(rows in the batch) against the current-version slice of the target, far cheaper than rebuilding a history table by hand.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;ETL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — etl&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Load-mode and incremental-materialization problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Transform&lt;/span&gt;
&lt;span&gt;Topic — data-transformation&lt;/span&gt;
&lt;strong&gt;Upsert, merge and SCD2 history problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-transformation" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Built-in &amp;amp; custom quality checks
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Quality checks live in the asset file and run right after it — a blocking failure stops the DAG
&lt;/h3&gt;

&lt;p&gt;The feature that sells Bruin to a data engineer who has been burned by silent data corruption is that &lt;strong&gt;tests are part of the asset, not a separate suite you hope someone wired up&lt;/strong&gt;. You declare column-level checks and SQL-based custom checks in the same &lt;code&gt;@bruin&lt;/code&gt; block as the query, and Bruin runs them immediately after the asset produces data. A failing check with &lt;code&gt;blocking: true&lt;/code&gt; fails the asset and prevents its downstream from running — the bad data never propagates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Built-in column checks.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Existence and uniqueness.&lt;/strong&gt; &lt;code&gt;not_null&lt;/code&gt; (no nulls in the column) and &lt;code&gt;unique&lt;/code&gt; (no value appears twice) — the two you attach to almost every key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sign and range.&lt;/strong&gt; &lt;code&gt;positive&lt;/code&gt;, &lt;code&gt;negative&lt;/code&gt;, &lt;code&gt;non_negative&lt;/code&gt;, and &lt;code&gt;min&lt;/code&gt; / &lt;code&gt;max&lt;/code&gt; with a &lt;code&gt;value&lt;/code&gt; threshold (numbers or dates) constrain the domain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Membership and shape.&lt;/strong&gt; &lt;code&gt;accepted_values&lt;/code&gt; with a &lt;code&gt;value&lt;/code&gt; list restricts to an enum; &lt;code&gt;pattern&lt;/code&gt; with a regex enforces a format.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Referential integrity.&lt;/strong&gt; &lt;code&gt;relationships&lt;/code&gt; verifies every non-null value exists in a parent column named by the column's &lt;code&gt;foreign_key&lt;/code&gt; metadata — a foreign-key test with no join written by you.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Custom checks — for logic the built-ins can't express.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A &lt;code&gt;custom_checks&lt;/code&gt; entry is a SQL query plus an expected &lt;code&gt;value&lt;/code&gt;.&lt;/strong&gt; Bruin runs the query and passes the check only if the returned integer equals &lt;code&gt;value&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Encode business invariants.&lt;/strong&gt; "Row count is greater than zero," "client X has exactly 15 credits for June," "no order has a negative net after refunds" — anything you can phrase as a query returning a number.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;blocking&lt;/code&gt; and &lt;code&gt;description&lt;/code&gt;.&lt;/strong&gt; Each check can set &lt;code&gt;blocking&lt;/code&gt; (whether a failure stops downstream) and a human &lt;code&gt;description&lt;/code&gt; for the report.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How checks run and gate the pipeline.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;After the asset, not before.&lt;/strong&gt; Checks execute once the asset has produced its table, then validate the produced data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;blocking&lt;/code&gt; defaults to &lt;code&gt;true&lt;/code&gt;.&lt;/strong&gt; A blocking failure marks the asset failed and stops its downstream; set &lt;code&gt;blocking: false&lt;/code&gt; to record the failure without blocking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;retries&lt;/code&gt;.&lt;/strong&gt; A check can retry on failure independently, resolving through the chain check → asset → pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run checks alone.&lt;/strong&gt; &lt;code&gt;bruin run --only checks assets/my_asset.sql&lt;/code&gt; re-validates without recomputing the table.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fupd57xeqtq9mtcb9pu6x.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fupd57xeqtq9mtcb9pu6x.jpeg" alt="Iconographic Bruin quality-checks diagram — column checks like not_null and unique attached to columns, a custom SQL check card with an expected value, a blocking gate that stops downstream assets on failure, and a passed/failed indicator." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — column checks on a keyed table
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday quality pattern is a handful of column checks that encode what "correct" means for a table: the key is unique and never null, a count is positive, a status is one of a fixed set. You attach them in the &lt;code&gt;columns&lt;/code&gt; block and Bruin runs them after the asset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Add checks to a &lt;code&gt;dataset.player_stats&lt;/code&gt; table so that &lt;code&gt;name&lt;/code&gt; is unique and not null, &lt;code&gt;player_count&lt;/code&gt; is not null and positive, and &lt;code&gt;status&lt;/code&gt; is one of &lt;code&gt;active&lt;/code&gt; or &lt;code&gt;inactive&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;columns&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;name&lt;/span&gt;        &lt;span class="c1"&gt;# unique, not null&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;player_count&lt;/span&gt;  &lt;span class="c1"&gt;# not null, positive&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;status&lt;/span&gt;      &lt;span class="c1"&gt;# accepted_values: active/inactive&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: dataset.player_stats
type: duckdb.sql
materialization:
  type: table
depends:
  - dataset.players

columns:
  - name: name
    type: string
    checks:
      - name: not_null
      - name: unique
  - name: player_count
    type: integer
    checks:
      - name: not_null
      - name: positive
  - name: status
    type: string
    checks:
      - name: accepted_values
        value: [active, inactive]
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;player_count&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'active'&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dataset&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;players&lt;/span&gt;
&lt;span class="k"&gt;group&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Bruin first runs the query and materializes &lt;code&gt;dataset.player_stats&lt;/code&gt;. Then it runs the checks column by column: &lt;code&gt;not_null&lt;/code&gt; and &lt;code&gt;unique&lt;/code&gt; on &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;not_null&lt;/code&gt; and &lt;code&gt;positive&lt;/code&gt; on &lt;code&gt;player_count&lt;/code&gt;, and &lt;code&gt;accepted_values&lt;/code&gt; on &lt;code&gt;status&lt;/code&gt;. Each check is a generated SQL query counting violations; zero violations passes. Because &lt;code&gt;blocking&lt;/code&gt; defaults to &lt;code&gt;true&lt;/code&gt;, any failing check fails the asset and stops anything that &lt;code&gt;depends&lt;/code&gt; on it — so a downstream dashboard model never reads a table with a duplicate key.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;column&lt;/th&gt;
&lt;th&gt;check&lt;/th&gt;
&lt;th&gt;passes when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;name&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;not_null&lt;/code&gt;, &lt;code&gt;unique&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;no nulls, no duplicates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;player_count&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;not_null&lt;/code&gt;, &lt;code&gt;positive&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;no nulls, all &lt;code&gt;&amp;gt; 0&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;status&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;accepted_values&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;every value in &lt;code&gt;{active, inactive}&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Attach &lt;code&gt;not_null&lt;/code&gt; + &lt;code&gt;unique&lt;/code&gt; to every key as a reflex; add &lt;code&gt;accepted_values&lt;/code&gt; and &lt;code&gt;positive&lt;/code&gt; wherever the domain is known — the checks cost one query each and catch the drift that silently corrupts downstream tables.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bruin interview question on business-rule validation
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A finance table &lt;code&gt;tier2.client_credits&lt;/code&gt; must satisfy a rule the built-in checks can't express: for June 2024, client X should have exactly 15 rows where &lt;code&gt;credits_spent = 1&lt;/code&gt;. If it doesn't, the run must fail before any downstream asset reads the table. How do you encode that in Bruin?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using a blocking custom check
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: tier2.client_credits
type: bq.sql
materialization:
  type: table

custom_checks:
  - name: Client X has 15 credits for June 2024
    description: Guards against the ACME-1234 miscount regression.
    value: 15
    blocking: true
    query: |
      SELECT count(*)
      FROM `tier2.client_credits`
      WHERE client = 'client_x'
        AND date_trunc(start_date, month) = '2024-06-01'
        AND credits_spent = 1
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;start_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;credits_spent&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;staging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;credits&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;stage&lt;/th&gt;
&lt;th&gt;what runs&lt;/th&gt;
&lt;th&gt;outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;asset query materializes &lt;code&gt;tier2.client_credits&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;table built&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;custom check query runs, returns a count&lt;/td&gt;
&lt;td&gt;e.g. &lt;code&gt;15&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Bruin compares count to &lt;code&gt;value: 15&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;equal → pass; not equal → fail&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;on fail with &lt;code&gt;blocking: true&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;asset marked failed, downstream stopped&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Bruin materializes the table first, then executes the &lt;code&gt;custom_checks&lt;/code&gt; query against the just-written data.&lt;/li&gt;
&lt;li&gt;The query returns a single integer — the number of matching credit rows for client X in June 2024.&lt;/li&gt;
&lt;li&gt;Bruin passes the check only if that integer equals the declared &lt;code&gt;value&lt;/code&gt; of &lt;code&gt;15&lt;/code&gt;; any other number fails.&lt;/li&gt;
&lt;li&gt;Because &lt;code&gt;blocking: true&lt;/code&gt;, a failure marks &lt;code&gt;tier2.client_credits&lt;/code&gt; as failed and prevents every asset that &lt;code&gt;depends&lt;/code&gt; on it from running, so the miscount never reaches a report.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;returned count&lt;/th&gt;
&lt;th&gt;check result&lt;/th&gt;
&lt;th&gt;downstream&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;14 or 16&lt;/td&gt;
&lt;td&gt;fail (blocking)&lt;/td&gt;
&lt;td&gt;stopped&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Custom check&lt;/strong&gt;&lt;/strong&gt; — an arbitrary SQL query plus an expected &lt;code&gt;value&lt;/code&gt; lets you encode any business invariant the built-in checks can't express.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Expected-value equality&lt;/strong&gt;&lt;/strong&gt; — the pass condition is "query returns exactly &lt;code&gt;value&lt;/code&gt;," which turns a fuzzy business rule into a deterministic gate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;blocking gate&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;blocking: true&lt;/code&gt; converts a failed check into a hard stop, so corrupt data cannot flow downstream — a failed run is recoverable, a silently wrong report is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Co-located with the asset&lt;/strong&gt;&lt;/strong&gt; — the check lives in the same file as the table it guards, so the invariant is versioned and reviewed alongside the query that must satisfy it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — each check is one aggregate query, O(rows scanned by the query); use partitioned predicates to keep the scan cheap on large tables.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Quality&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — data-quality&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Built-in and custom data-quality-check problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-quality" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Validation&lt;/span&gt;
&lt;span&gt;Topic — data-validation&lt;/span&gt;
&lt;strong&gt;Referential-integrity and constraint-validation problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-validation" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Ingestion with ingestr &amp;amp; multi-destination DAGs
&lt;/h2&gt;
&lt;h3&gt;
  
  
  An ingestr asset lands raw data, SQL and Python transform it, and bruin run executes the whole DAG
&lt;/h3&gt;

&lt;p&gt;The last piece that makes Bruin end-to-end is &lt;strong&gt;ingestion&lt;/strong&gt;: instead of a separate connector platform, you declare an &lt;code&gt;ingestr&lt;/code&gt; asset — a small YAML file — and Bruin pulls data from a source into a destination as the first node of the same DAG that transforms and tests it. &lt;code&gt;ingestr&lt;/code&gt; is Bruin's open-source ingestion CLI, and one line of config replaces a hand-written extract job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The ingestr asset.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;type: ingestr&lt;/code&gt; in a &lt;code&gt;.asset.yml&lt;/code&gt;.&lt;/strong&gt; No code body — just &lt;code&gt;parameters&lt;/code&gt; naming the source and destination.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;parameters&lt;/code&gt;.&lt;/strong&gt; &lt;code&gt;source_connection&lt;/code&gt; and &lt;code&gt;source_table&lt;/code&gt; say where data comes from; &lt;code&gt;destination&lt;/code&gt; says where it lands. The connections themselves live in &lt;code&gt;.bruin.yml&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Any source to any destination.&lt;/strong&gt; ingestr covers databases, SaaS APIs, and file sources, so "there is no connector for this" is rarely the blocker it is with a hand-rolled script.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The runner and its flags.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;bruin run&lt;/code&gt;.&lt;/strong&gt; Execute the whole pipeline — every asset, in dependency order.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;bruin run assets/x.sql&lt;/code&gt;.&lt;/strong&gt; Run a single asset (and, with &lt;code&gt;--downstream&lt;/code&gt;, everything that depends on it).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;bruin run --tag daily&lt;/code&gt;.&lt;/strong&gt; Run only assets carrying a tag — the same graph, sliced by schedule or domain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;bruin run --full-refresh&lt;/code&gt;.&lt;/strong&gt; Drop and rebuild instead of incrementally updating, for a clean reload.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;bruin run --start-date … --end-date …&lt;/code&gt;.&lt;/strong&gt; Bound the window for &lt;code&gt;time_interval&lt;/code&gt; and ingestr assets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;bruin validate&lt;/code&gt;.&lt;/strong&gt; Statically check every asset — parse the metadata, verify types and dependencies — without running anything, so CI catches a broken pipeline before it touches data.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The DAG and multi-destination pipelines.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;depends&lt;/code&gt; builds the graph.&lt;/strong&gt; Bruin reads every asset's &lt;code&gt;depends&lt;/code&gt;, assembles a DAG, and runs it in topological order — ingest first, then staging, then marts, then checks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One pipeline can span platforms.&lt;/strong&gt; An ingestr asset can land data in Postgres, a &lt;code&gt;duckdb.sql&lt;/code&gt; asset can transform it, and a &lt;code&gt;bq.sql&lt;/code&gt; asset can publish it — Bruin routes each asset to its own connection by &lt;code&gt;type&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No external orchestrator for a single pipeline.&lt;/strong&gt; &lt;code&gt;bruin run&lt;/code&gt; &lt;em&gt;is&lt;/em&gt; the scheduler for one graph; when you need cluster-scale scheduling, Bruin can deploy the same assets to Airflow.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsj89b5aitemr240dpu7b.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsj89b5aitemr240dpu7b.jpeg" alt="Iconographic Bruin ingestr diagram — external sources flowing through an ingestr asset into a destination, then SQL and Python assets transforming the data across multiple platforms, all executed as one topologically ordered DAG by bruin run." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — an ingestr asset feeding a SQL transform
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday ingestion pattern is a two-asset chain: an &lt;code&gt;ingestr&lt;/code&gt; asset lands a raw table, and a SQL asset depends on it to build a clean model. Running the pipeline executes them in order without you sequencing anything by hand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Ingest a &lt;code&gt;profiles&lt;/code&gt; table from a &lt;code&gt;chess&lt;/code&gt; source into DuckDB, then build a &lt;code&gt;dataset.player_stats&lt;/code&gt; aggregate that depends on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; Two files under &lt;code&gt;assets/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;dataset.players&lt;/span&gt;
&lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ingestr&lt;/span&gt;
&lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;destination&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;duckdb&lt;/span&gt;
  &lt;span class="na"&gt;source_connection&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;chess-default&lt;/span&gt;
  &lt;span class="na"&gt;source_table&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;profiles&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: dataset.player_stats
type: duckdb.sql
materialization:
  type: table
depends:
  - dataset.players
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;player_count&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dataset&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;players&lt;/span&gt;
&lt;span class="k"&gt;group&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The &lt;code&gt;.asset.yml&lt;/code&gt; file declares an &lt;code&gt;ingestr&lt;/code&gt; asset: it reads the &lt;code&gt;chess-default&lt;/code&gt; connection's &lt;code&gt;profiles&lt;/code&gt; table and lands it in DuckDB as &lt;code&gt;dataset.players&lt;/code&gt;. The SQL asset lists &lt;code&gt;dataset.players&lt;/code&gt; under &lt;code&gt;depends&lt;/code&gt;, so Bruin knows the ingestion must finish before the aggregate runs. &lt;code&gt;bruin run&lt;/code&gt; parses both, orders them ingest → transform, executes ingestr to load the raw table, then runs the &lt;code&gt;SELECT&lt;/code&gt; to build &lt;code&gt;dataset.player_stats&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;asset&lt;/th&gt;
&lt;th&gt;type&lt;/th&gt;
&lt;th&gt;produces&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;dataset.players&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ingestr&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;raw &lt;code&gt;profiles&lt;/code&gt; landed in DuckDB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;dataset.player_stats&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;duckdb.sql&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;aggregated counts (runs second)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Model ingestion as an asset, not a pre-step: put the source in an &lt;code&gt;ingestr&lt;/code&gt; &lt;code&gt;.asset.yml&lt;/code&gt; and let every downstream model &lt;code&gt;depends&lt;/code&gt; on it, so the whole extract-transform chain runs and validates in one &lt;code&gt;bruin run&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bruin interview question on orchestrating a mixed pipeline
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You have four assets — an ingestr load, a SQL staging model, a Python enrichment job, and a final SQL mart with quality checks. An interviewer asks: how does Bruin run these in the correct order across two platforms in one command, and how would you verify the graph is valid before it touches production data?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using depends-driven topological execution and bruin validate
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: mart.daily_active
type: bq.sql
materialization:
  type: table
  strategy: merge
depends:
  - raw.events          # ingestr asset
  - staging.sessions    # duckdb.sql asset
  - features.user_score # python asset

columns:
  - name: user_id
    type: integer
    primary_key: true
    checks:
      - name: not_null
      - name: unique
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;active_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;staging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sessions&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
&lt;span class="k"&gt;join&lt;/span&gt; &lt;span class="n"&gt;features&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;user_score&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt; &lt;span class="k"&gt;using&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;asset&lt;/th&gt;
&lt;th&gt;type&lt;/th&gt;
&lt;th&gt;depends on&lt;/th&gt;
&lt;th&gt;run position&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;raw.events&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;ingestr&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;staging.sessions&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;duckdb.sql&lt;/td&gt;
&lt;td&gt;&lt;code&gt;raw.events&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;features.user_score&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;python&lt;/td&gt;
&lt;td&gt;&lt;code&gt;staging.sessions&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mart.daily_active&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;bq.sql&lt;/td&gt;
&lt;td&gt;all three&lt;/td&gt;
&lt;td&gt;4, then checks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Bruin parses all four assets and builds the DAG from their &lt;code&gt;depends&lt;/code&gt; edges: ingest → staging → features → mart.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;bruin validate&lt;/code&gt; first runs statically — it parses every &lt;code&gt;@bruin&lt;/code&gt; block, checks that referenced dependencies exist, and verifies types and column definitions, all &lt;strong&gt;without executing&lt;/strong&gt; a single query.&lt;/li&gt;
&lt;li&gt;On &lt;code&gt;bruin run&lt;/code&gt;, Bruin executes in topological order, routing each asset to its platform by &lt;code&gt;type&lt;/code&gt; (ingestr, DuckDB, Python, BigQuery) — one command, two-plus platforms.&lt;/li&gt;
&lt;li&gt;After &lt;code&gt;mart.daily_active&lt;/code&gt; materializes via &lt;code&gt;merge&lt;/code&gt;, its &lt;code&gt;not_null&lt;/code&gt; and &lt;code&gt;unique&lt;/code&gt; checks on &lt;code&gt;user_id&lt;/code&gt; run; a blocking failure would stop anything downstream of the mart.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;command&lt;/th&gt;
&lt;th&gt;effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;pre-flight&lt;/td&gt;
&lt;td&gt;&lt;code&gt;bruin validate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;graph + metadata verified, nothing run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;execute&lt;/td&gt;
&lt;td&gt;&lt;code&gt;bruin run&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;four assets in order, then checks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;depends-driven DAG&lt;/strong&gt;&lt;/strong&gt; — the union of every asset's &lt;code&gt;depends&lt;/code&gt; list is the graph; Bruin topologically sorts it, so order is derived, never hand-maintained.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cross-platform routing&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;type&lt;/code&gt; sends each asset to its own connection, so a single DAG spans ingestr, DuckDB, Python, and BigQuery in one run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;validate before run&lt;/strong&gt;&lt;/strong&gt; — static validation catches missing dependencies and metadata errors in CI, before any query touches production data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Checks as the final gate&lt;/strong&gt;&lt;/strong&gt; — quality checks on the mart run after materialization and block downstream on failure, closing the loop from ingest to validated output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the DAG sort is O(V + E); total runtime is dominated by the assets themselves, and &lt;code&gt;bruin run&lt;/code&gt; adds no orchestrator overhead for a single pipeline.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;ETL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — etl&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Ingestion and extract-load pipeline problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Pipelines&lt;/span&gt;
&lt;span&gt;Topic — pipelines&lt;/span&gt;
&lt;strong&gt;DAG-orchestration and multi-platform pipeline problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — Bruin recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Minimal SQL asset.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: dashboard.hello
type: duckdb.sql
materialization:
  type: table
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Python asset with a dependency.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;@bruin
name: analytics.features
type: python
depends:
  - staging.players
@bruin&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;build features here&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;ingestr asset (any source to any destination).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;raw.profiles&lt;/span&gt;
&lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ingestr&lt;/span&gt;
&lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;destination&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;bigquery&lt;/span&gt;
  &lt;span class="na"&gt;source_connection&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;chess-default&lt;/span&gt;
  &lt;span class="na"&gt;source_table&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;profiles&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Merge / upsert materialization.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: dim.customers
type: bq.sql
materialization:
  type: table
  strategy: merge
columns:
  - name: id
    type: integer
    primary_key: true
  - name: email
    type: string
    update_on_merge: true
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;staging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customers&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Column + custom quality checks.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* @bruin
name: dataset.orders
type: duckdb.sql
materialization:
  type: table
columns:
  - name: order_id
    type: integer
    checks:
      - name: not_null
      - name: unique
custom_checks:
  - name: table is not empty
    query: SELECT count(*) FROM dataset.orders
    value: 1
    blocking: true
@bruin */&lt;/span&gt;

&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;order_id&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;staging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Common &lt;code&gt;bruin&lt;/code&gt; commands.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;bruin init default my-pipeline      &lt;span class="c"&gt;# scaffold a project&lt;/span&gt;
bruin validate                      &lt;span class="c"&gt;# static-check every asset, run nothing&lt;/span&gt;
bruin run                           &lt;span class="c"&gt;# run the whole DAG in order&lt;/span&gt;
bruin run assets/orders.sql &lt;span class="nt"&gt;--downstream&lt;/span&gt;
bruin run &lt;span class="nt"&gt;--tag&lt;/span&gt; daily &lt;span class="nt"&gt;--full-refresh&lt;/span&gt;
bruin run &lt;span class="nt"&gt;--only&lt;/span&gt; checks assets/orders.sql
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Materialization picker.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Small, re-pullable table&lt;/td&gt;
&lt;td&gt;&lt;code&gt;create+replace&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Append-only events&lt;/td&gt;
&lt;td&gt;&lt;code&gt;append&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mutable rows keyed by an id&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;merge&lt;/code&gt; + &lt;code&gt;primary_key&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reprocess a partition&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;delete+insert&lt;/code&gt; + &lt;code&gt;incremental_key&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full change history required&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;scd2_by_column&lt;/code&gt; / &lt;code&gt;scd2_by_time&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is Bruin?
&lt;/h3&gt;

&lt;p&gt;Bruin is an open-source data pipeline framework — a single Go CLI plus a VS Code extension — that runs SQL transforms, Python jobs, ingestion, and data-quality checks from one project. You define &lt;em&gt;assets&lt;/em&gt; (SQL, Python, or YAML files) whose metadata lives in an inline &lt;code&gt;@bruin&lt;/code&gt; comment block, declare dependencies with &lt;code&gt;depends&lt;/code&gt;, and run the whole DAG with &lt;code&gt;bruin run&lt;/code&gt;. It is a tool you install and version in your repo, not a hosted platform.&lt;/p&gt;

&lt;h3&gt;
  
  
  How is Bruin different from dbt?
&lt;/h3&gt;

&lt;p&gt;dbt is transform-only: it compiles and runs SQL (and Python) models but leaves ingestion, testing infrastructure, and scheduling to other tools. Bruin covers the same SQL-modelling ground and adds native Python assets, &lt;code&gt;ingestr&lt;/code&gt;-based ingestion, built-in column and custom quality checks, and a runner that executes the DAG without a separate orchestrator. In short, Bruin aims to be SQL + Python + ingestion + quality in one framework, where dbt is the transform layer of a larger stack.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do Bruin quality checks work?
&lt;/h3&gt;

&lt;p&gt;You declare checks in the same &lt;code&gt;@bruin&lt;/code&gt; block as the asset. Column checks like &lt;code&gt;not_null&lt;/code&gt;, &lt;code&gt;unique&lt;/code&gt;, &lt;code&gt;positive&lt;/code&gt;, &lt;code&gt;accepted_values&lt;/code&gt;, &lt;code&gt;pattern&lt;/code&gt;, &lt;code&gt;min&lt;/code&gt;/&lt;code&gt;max&lt;/code&gt;, and &lt;code&gt;relationships&lt;/code&gt; attach to a column; &lt;code&gt;custom_checks&lt;/code&gt; are SQL queries with an expected integer &lt;code&gt;value&lt;/code&gt;. Checks run right after the asset materializes, and a failing check with &lt;code&gt;blocking: true&lt;/code&gt; (the default) fails the asset and stops its downstream, so bad data never propagates. You can re-run just the checks with &lt;code&gt;bruin run --only checks&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  What materialization strategies does Bruin support?
&lt;/h3&gt;

&lt;p&gt;For &lt;code&gt;type: table&lt;/code&gt;, Bruin supports &lt;code&gt;create+replace&lt;/code&gt; (full overwrite, the default), &lt;code&gt;append&lt;/code&gt; (immutable insert), &lt;code&gt;delete+insert&lt;/code&gt; and &lt;code&gt;time_interval&lt;/code&gt; (incremental with an &lt;code&gt;incremental_key&lt;/code&gt;), &lt;code&gt;truncate+insert&lt;/code&gt; (replace data, keep the table), &lt;code&gt;merge&lt;/code&gt; (upsert on a &lt;code&gt;primary_key&lt;/code&gt;), &lt;code&gt;ddl&lt;/code&gt; (create an empty typed table), and &lt;code&gt;scd2_by_column&lt;/code&gt; / &lt;code&gt;scd2_by_time&lt;/code&gt; for full history with &lt;code&gt;_valid_from&lt;/code&gt; / &lt;code&gt;_valid_until&lt;/code&gt; / &lt;code&gt;_is_current&lt;/code&gt; columns. You write a plain &lt;code&gt;SELECT&lt;/code&gt; and Bruin generates the DDL and DML for the strategy you chose. &lt;code&gt;type: view&lt;/code&gt; persists the query as a view instead.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is ingestr and how does Bruin use it?
&lt;/h3&gt;

&lt;p&gt;ingestr is Bruin's open-source ingestion CLI that copies data from any supported source (databases, SaaS APIs, files) into any supported destination. In Bruin you use it by declaring an &lt;code&gt;ingestr&lt;/code&gt; asset — a &lt;code&gt;.asset.yml&lt;/code&gt; file with &lt;code&gt;parameters&lt;/code&gt; naming the &lt;code&gt;source_connection&lt;/code&gt;, &lt;code&gt;source_table&lt;/code&gt;, and &lt;code&gt;destination&lt;/code&gt;, with the connection secrets kept in &lt;code&gt;.bruin.yml&lt;/code&gt;. That asset becomes the first node of your DAG, so ingestion, transformation, and quality checks all run in the same &lt;code&gt;bruin run&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I need an orchestrator to run Bruin?
&lt;/h3&gt;

&lt;p&gt;Not for a single pipeline. &lt;code&gt;bruin run&lt;/code&gt; reads the &lt;code&gt;depends&lt;/code&gt; edges, builds the DAG, and executes assets in topological order itself, so it is the scheduler for one graph — you can run it locally, in CI, or in a container. When you need cluster-scale scheduling across many pipelines, Bruin can deploy the same assets to Airflow, and settings like &lt;code&gt;rerun_cooldown&lt;/code&gt; translate to Airflow retry semantics. Use &lt;code&gt;bruin validate&lt;/code&gt; in CI to catch a broken graph before it runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every Bruin idea above, from the depends-driven DAG to the merge materialization and the blocking custom check, maps to a hands-on practice room where you build the pipeline against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you gate this pipeline on data quality?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice pipeline problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/data-quality" rel="noopener noreferrer"&gt;Data-quality drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>A Local Lakehouse on Your Laptop: DuckDB + Iceberg + Trino for Zero-Cloud Dev</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Sun, 27 Sep 2026 17:39:50 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/a-local-lakehouse-on-your-laptop-duckdb-iceberg-trino-for-zero-cloud-dev-226a</link>
      <guid>https://dev.to/gowthampotureddi/a-local-lakehouse-on-your-laptop-duckdb-iceberg-trino-for-zero-cloud-dev-226a</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;local lakehouse&lt;/code&gt;&lt;/strong&gt; means the whole modern data stack — an open table format, object storage, and one or more query engines — running on your machine with nothing rented from a cloud provider. The pieces that used to require an S3 bucket, a Glue catalog, and a Spark cluster now install as a Python wheel, a single binary, and a Docker container. You land Parquet on your own disk, register it as an Apache Iceberg table in a catalog you host, and query it from DuckDB in the same process or from Trino over a socket — all offline, all free, all deterministic.&lt;/p&gt;

&lt;p&gt;That shape matters because the slowest part of data engineering is rarely the query; it is the loop. Every change you make to a transform, a partition spec, or a merge key normally has to round-trip through a shared dev warehouse: push, wait for the orchestrator, burn warehouse credits, discover the schema drifted, repeat. A lakehouse you can spin up on a laptop collapses that loop to seconds and makes it reproducible in CI. This guide is a hands-on walkthrough of the four moving parts an interviewer — or your own future test suite — will actually probe: DuckDB reading and writing Parquet, a local Iceberg catalog on the filesystem or MinIO, Trino querying the exact same tables, and partitioning plus compaction done locally. Each section pairs the concepts with a Solution-Tail interview answer: code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffe9gxo30nsprqlqy7g57.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffe9gxo30nsprqlqy7g57.jpeg" alt="PipeCode blog header for a local lakehouse — bold white headline 'Local Lakehouse' with subtitle 'DuckDB · Iceberg · Trino · zero cloud' and a stylised laptop-as-warehouse scene on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;ETL practice library →&lt;/a&gt;, tune scans and file layout on the &lt;a href="https://pipecode.ai/explore/practice/topic/optimization" rel="noopener noreferrer"&gt;optimization practice set →&lt;/a&gt;, and rehearse table layout on the &lt;a href="https://pipecode.ai/explore/practice/topic/partitioning" rel="noopener noreferrer"&gt;partitioning practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why a local lakehouse beats cloud dev loops in 2026&lt;/li&gt;
&lt;li&gt;DuckDB: reading and writing Parquet on the local filesystem&lt;/li&gt;
&lt;li&gt;A local Iceberg catalog on the filesystem or MinIO&lt;/li&gt;
&lt;li&gt;Trino querying the same Iceberg tables&lt;/li&gt;
&lt;li&gt;Partitioning and compaction on your laptop&lt;/li&gt;
&lt;li&gt;Cheat sheet — local lakehouse recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why a local lakehouse beats cloud dev loops in 2026
&lt;/h2&gt;

&lt;h3&gt;
  
  
  A lakehouse is three swappable layers — table format, storage, engine — and every one of them now runs on a laptop
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;a lakehouse is an open table format over plain files, plus a catalog that names those files as tables, plus any engine that speaks the format — and none of those three layers requires a cloud account to run.&lt;/strong&gt; Once you internalise that separation, "run it locally" stops being a hack and becomes the obvious default for development.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The three layers, decoupled.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Table format (Apache Iceberg).&lt;/strong&gt; Iceberg is a specification, not a service: a table is a tree of JSON and Avro metadata files (&lt;code&gt;metadata.json&lt;/code&gt; → manifest list → manifests) that point at immutable Parquet data files. Nothing about that tree needs the cloud; it can sit in any directory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage.&lt;/strong&gt; The data and metadata files live somewhere with a path. That "somewhere" can be your local filesystem (&lt;code&gt;file:///Users/you/warehouse&lt;/code&gt;) or a local S3-compatible server like MinIO (&lt;code&gt;s3://warehouse&lt;/code&gt;). The format does not care which.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Engine.&lt;/strong&gt; DuckDB, Trino, Spark, and pyiceberg are all clients that read the same Iceberg metadata and Parquet files. You pick the engine per task — DuckDB for a fast in-process scan, Trino for distributed SQL, pyiceberg for writes — and they agree because they share one catalog.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The zero-cloud loop.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Land Parquet on disk.&lt;/strong&gt; A generator, a &lt;code&gt;pandas&lt;/code&gt;/&lt;code&gt;pyarrow&lt;/code&gt; frame, or a &lt;code&gt;COPY&lt;/code&gt; statement writes columnar files to a folder. No upload, no credentials.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Register and query in-process.&lt;/strong&gt; DuckDB reads those files with zero server; pyiceberg or Trino registers them as an Iceberg table so schema, snapshots, and partitioning become first-class.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Iterate at memory speed.&lt;/strong&gt; Because everything is local, a full write-query-verify cycle is milliseconds to seconds, not the minutes a shared warehouse round-trip costs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why interviewers and CI both care.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Determinism.&lt;/strong&gt; A fixed local dataset plus Iceberg snapshot isolation means a test reads exactly the rows it expects, every run, with no "someone else's job changed the shared table" flakiness.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offline and free.&lt;/strong&gt; CI runners need no cloud secrets, no network egress, and no per-query bill; the lakehouse spins up inside the job and is torn down with it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fidelity.&lt;/strong&gt; Because the &lt;em&gt;format&lt;/em&gt; is identical to production Iceberg, a local test exercises the real partition spec, the real merge semantics, and the real compaction — not a sqlite stand-in that behaves differently.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What an interviewer listens for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you separate &lt;strong&gt;"table format vs storage vs engine"&lt;/strong&gt; cleanly, rather than treating "the lakehouse" as one product? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you place &lt;strong&gt;DuckDB as an in-process reader and Trino as the distributed engine&lt;/strong&gt;, using the same tables? — required framing.&lt;/li&gt;
&lt;li&gt;Do you reach for a &lt;strong&gt;local Iceberg catalog for tests&lt;/strong&gt; instead of mocking the warehouse? — the whole point of this post.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — the whole loop in two DuckDB statements
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; Before any catalog or engine setup, the smallest possible proof that a lakehouse loop needs no cloud is a single DuckDB session that writes Parquet to local disk and reads it straight back. DuckDB runs in-process — there is no server to start, no port to bind, no credentials — so this is the innermost loop every later layer wraps.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Using only DuckDB, write two order rows to a local Parquet file and read them back as an aggregate, with no server and no cloud storage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;amount&lt;/th&gt;
&lt;th&gt;order_date&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;42.50&lt;/td&gt;
&lt;td&gt;2026-01-05&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;linus&lt;/td&gt;
&lt;td&gt;17.00&lt;/td&gt;
&lt;td&gt;2026-02-11&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- duckdb (in-process; `duckdb` CLI or the Python module — no server)&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;VALUES&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'ada'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-05'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'linus'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;17&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="s1"&gt;'2026-02-11'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;order_date&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;TO&lt;/span&gt; &lt;span class="s1"&gt;'lake/orders.parquet'&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;FORMAT&lt;/span&gt; &lt;span class="n"&gt;PARQUET&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="s1"&gt;'lake/orders.parquet'&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;customer&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;COPY (...) TO 'lake/orders.parquet' (FORMAT PARQUET)&lt;/code&gt; materialises the query result as a single columnar Parquet file on your local filesystem — DuckDB creates the &lt;code&gt;lake/&lt;/code&gt; directory content and writes typed columns with statistics. The second statement reads that file directly by naming its path in the &lt;code&gt;FROM&lt;/code&gt; clause; DuckDB's Parquet reader opens the file, applies the &lt;code&gt;GROUP BY&lt;/code&gt;, and returns rows. There is no &lt;code&gt;CREATE TABLE&lt;/code&gt;, no connection string, and no object store — the file &lt;em&gt;is&lt;/em&gt; the table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;42.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;linus&lt;/td&gt;
&lt;td&gt;17.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If a step in your pipeline can be expressed as "read Parquet, transform, write Parquet," it can run locally with zero cloud — the catalog and Trino layers you add later only make those files behave like managed tables.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. DuckDB: reading and writing Parquet on the local filesystem
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;read_parquet&lt;/code&gt;, &lt;code&gt;COPY ... PARTITION_BY&lt;/code&gt;, and the &lt;code&gt;iceberg&lt;/code&gt; extension make DuckDB the fast local reader and Parquet producer
&lt;/h3&gt;

&lt;p&gt;DuckDB is the engine you reach for first in a local lakehouse because it is an embedded, columnar SQL engine with a first-class Parquet reader and writer and an optional &lt;code&gt;iceberg&lt;/code&gt; extension. It has no server, starts in milliseconds, and pushes filters and projections down into Parquet so a scan touches only the bytes it needs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reading Parquet.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Glob paths.&lt;/strong&gt; &lt;code&gt;read_parquet('lake/orders/**/*.parquet')&lt;/code&gt; reads a whole tree; you can also just put the path string in &lt;code&gt;FROM&lt;/code&gt;. DuckDB unifies the files into one relation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Projection pushdown.&lt;/strong&gt; Selecting three columns from a hundred-column file reads only those column chunks — Parquet is columnar, and DuckDB honours it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Predicate pushdown + statistics.&lt;/strong&gt; &lt;code&gt;WHERE amount &amp;gt; 100&lt;/code&gt; uses each row group's min/max stats to skip groups that cannot match, so a filtered scan reads a fraction of the file.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Writing Parquet.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Single file.&lt;/strong&gt; &lt;code&gt;COPY tbl TO 'out.parquet' (FORMAT PARQUET)&lt;/code&gt; writes one file with a chosen compression (&lt;code&gt;CODEC 'zstd'&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hive-partitioned dataset.&lt;/strong&gt; &lt;code&gt;COPY tbl TO 'lake/orders' (FORMAT PARQUET, PARTITION_BY (event_date))&lt;/code&gt; writes &lt;code&gt;lake/orders/event_date=2026-02-11/data_0.parquet&lt;/code&gt;, encoding the partition value in the directory name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Row-group sizing.&lt;/strong&gt; &lt;code&gt;ROW_GROUP_SIZE&lt;/code&gt; controls how many rows per group, which trades scan granularity against file overhead.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The &lt;code&gt;iceberg&lt;/code&gt; extension.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Install once.&lt;/strong&gt; &lt;code&gt;INSTALL iceberg; LOAD iceberg;&lt;/code&gt; adds Iceberg support to a DuckDB session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scan a table.&lt;/strong&gt; &lt;code&gt;SELECT * FROM iceberg_scan('lake/warehouse/lake.db/orders', allow_moved_paths =&amp;gt; true)&lt;/code&gt; reads an Iceberg table by pointing at its metadata directory, resolving the current snapshot and its Parquet data files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attach a REST catalog.&lt;/strong&gt; Recent DuckDB can &lt;code&gt;ATTACH '' AS ice (TYPE ICEBERG, ENDPOINT 'http://localhost:8181/catalog')&lt;/code&gt; and then &lt;code&gt;SELECT * FROM ice.lake.orders&lt;/code&gt;, so DuckDB reads the same catalog Trino uses. Iceberg &lt;em&gt;writes&lt;/em&gt; from DuckDB are still newer/preview, so most local setups author tables with pyiceberg or Trino and use DuckDB as the fast reader.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy50xd8ngfqwmfrk3pktz.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy50xd8ngfqwmfrk3pktz.jpeg" alt="Iconographic DuckDB + Parquet diagram — an in-process DuckDB engine reading a glob of Parquet files with projection and predicate pushdown, writing a Hive-partitioned dataset, and reading an Iceberg table via the iceberg extension." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — write a partitioned Parquet dataset, then read one partition
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday DuckDB pattern is to write a Hive-partitioned Parquet directory and then read back only the partitions you need. Because the partition value is encoded in the folder path, a filter on that column lets DuckDB skip whole directories without opening their files — the local equivalent of partition pruning in a cloud warehouse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Write three orders partitioned by &lt;code&gt;event_date&lt;/code&gt;, then read back only &lt;code&gt;2026-02-11&lt;/code&gt; and show that DuckDB pruned the other partition.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;amount&lt;/th&gt;
&lt;th&gt;event_date&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;42.5&lt;/td&gt;
&lt;td&gt;2026-01-05&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;17.0&lt;/td&gt;
&lt;td&gt;2026-02-11&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;9.9&lt;/td&gt;
&lt;td&gt;2026-02-11&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- write a Hive-partitioned dataset&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;VALUES&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-05'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;11&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;17&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="s1"&gt;'2026-02-11'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="mi"&gt;9&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;9&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="s1"&gt;'2026-02-11'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_date&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;TO&lt;/span&gt; &lt;span class="s1"&gt;'lake/orders'&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;FORMAT&lt;/span&gt; &lt;span class="n"&gt;PARQUET&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;PARTITION_BY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event_date&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;

&lt;span class="c1"&gt;-- read back only one partition&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;revenue&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;read_parquet&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'lake/orders/**/*.parquet'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hive_partitioning&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;event_date&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="s1"&gt;'2026-02-11'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The &lt;code&gt;COPY ... PARTITION_BY (event_date)&lt;/code&gt; writes two directories — &lt;code&gt;event_date=2026-01-05/&lt;/code&gt; and &lt;code&gt;event_date=2026-02-11/&lt;/code&gt; — each holding the rows for that date. The read uses &lt;code&gt;hive_partitioning =&amp;gt; true&lt;/code&gt;, so DuckDB reconstructs &lt;code&gt;event_date&lt;/code&gt; from the directory names as a real column. The &lt;code&gt;WHERE event_date = DATE '2026-02-11'&lt;/code&gt; is evaluated against those directory values first, so DuckDB never opens the January file — it prunes the partition before any Parquet is read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;rows&lt;/th&gt;
&lt;th&gt;revenue&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;26.9&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Partition by the column you filter on most (usually a date), read with &lt;code&gt;hive_partitioning =&amp;gt; true&lt;/code&gt;, and DuckDB turns a directory layout into free partition pruning — the same idea Iceberg formalises with hidden partitioning.&lt;/p&gt;

&lt;h3&gt;
  
  
  DuckDB interview question on Parquet pushdown
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You have a directory &lt;code&gt;lake/events&lt;/code&gt; of Parquet files partitioned by &lt;code&gt;event_date&lt;/code&gt;, one hundred columns wide and billions of rows. An analyst needs &lt;code&gt;sum(amount)&lt;/code&gt; for a single month, &lt;code&gt;2026-02&lt;/code&gt;. Write the DuckDB query and explain, concretely, which bytes DuckDB avoids reading and why.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using hive partitioning with projection and predicate pushdown
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;INSTALL&lt;/span&gt; &lt;span class="n"&gt;parquet&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;-- bundled; explicit for clarity&lt;/span&gt;
&lt;span class="k"&gt;LOAD&lt;/span&gt; &lt;span class="n"&gt;parquet&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;feb_revenue&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;read_parquet&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'lake/events/**/*.parquet'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hive_partitioning&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;event_date&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="s1"&gt;'2026-02-01'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;event_date&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt;  &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="s1"&gt;'2026-03-01'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;stage&lt;/th&gt;
&lt;th&gt;what DuckDB inspects&lt;/th&gt;
&lt;th&gt;what it skips&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1. partition prune&lt;/td&gt;
&lt;td&gt;directory names &lt;code&gt;event_date=...&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;every folder outside Feb 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2. row-group stats&lt;/td&gt;
&lt;td&gt;Feb files' min/max for &lt;code&gt;event_date&lt;/code&gt;, &lt;code&gt;amount&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;row groups whose range misses the filter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3. projection&lt;/td&gt;
&lt;td&gt;only the &lt;code&gt;amount&lt;/code&gt; (and partition) column chunk&lt;/td&gt;
&lt;td&gt;the other 98 columns' chunks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4. scan&lt;/td&gt;
&lt;td&gt;matching &lt;code&gt;amount&lt;/code&gt; chunks in surviving row groups&lt;/td&gt;
&lt;td&gt;everything else on disk&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;DuckDB resolves &lt;code&gt;hive_partitioning =&amp;gt; true&lt;/code&gt; first, reading &lt;code&gt;event_date&lt;/code&gt; from the folder path, so the month predicate eliminates all non-February directories before a single Parquet byte is opened.&lt;/li&gt;
&lt;li&gt;Inside the surviving files, DuckDB reads each row group's footer statistics and drops groups whose &lt;code&gt;event_date&lt;/code&gt; min/max cannot satisfy the range.&lt;/li&gt;
&lt;li&gt;Because the query only references &lt;code&gt;amount&lt;/code&gt;, DuckDB's projection pushdown reads just that column's chunks — the other columns' bytes are never fetched from disk.&lt;/li&gt;
&lt;li&gt;Only the qualifying &lt;code&gt;amount&lt;/code&gt; chunks are decompressed and summed, so the scan cost is proportional to one month of one column, not the whole dataset.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;feb_revenue&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;128934.50&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Partition pruning&lt;/strong&gt;&lt;/strong&gt; — encoding &lt;code&gt;event_date&lt;/code&gt; in the directory path lets DuckDB eliminate whole folders from the plan using metadata alone, turning a full scan into a targeted one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Row-group statistics&lt;/strong&gt;&lt;/strong&gt; — Parquet stores per-group min/max, so even within a partition DuckDB skips groups that cannot match the predicate; this is why sorted or clustered data scans faster.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Projection pushdown&lt;/strong&gt;&lt;/strong&gt; — columnar storage means selecting one column reads one column's bytes; a hundred-column table costs the same as a one-column table for this query.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;In-process execution&lt;/strong&gt;&lt;/strong&gt; — no network hop to a storage service means the only cost is local disk I/O for the surviving bytes, which is what makes the loop feel instant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — I/O is O(qualifying row groups × selected columns), typically a tiny fraction of the dataset, versus O(all files × all columns) for a naive full read.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;ETL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — etl&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Local extract-and-load Parquet pipeline problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Pipelines&lt;/span&gt;
&lt;span&gt;Topic — pipelines&lt;/span&gt;
&lt;strong&gt;Pipeline-design and file-layout problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. A local Iceberg catalog on the filesystem or MinIO
&lt;/h2&gt;
&lt;h3&gt;
  
  
  The catalog is what turns a folder of Parquet into a table — SQLite for pure-filesystem, a REST catalog for sharing
&lt;/h3&gt;

&lt;p&gt;A pile of Parquet files is not a table; it becomes one only when a &lt;strong&gt;catalog&lt;/strong&gt; records the current metadata pointer, the schema, the snapshots, and the partition spec. The catalog is the single piece that makes DuckDB, Trino, and pyiceberg agree on what "the orders table" is. Locally you have two good choices, and picking between them is the decision an interviewer probes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What a catalog actually stores.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The metadata pointer.&lt;/strong&gt; For each table, the catalog holds the path to the current &lt;code&gt;metadata.json&lt;/code&gt;. That file references the manifest list, which references manifests, which reference the Parquet data files — the whole Iceberg tree hangs off the pointer the catalog owns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Atomic commits.&lt;/strong&gt; A write produces a new &lt;code&gt;metadata.json&lt;/code&gt;; the commit is the catalog atomically swapping the pointer from the old file to the new one. That swap is why Iceberg gives snapshot isolation — readers see the old snapshot until the pointer moves.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Namespaces and identity.&lt;/strong&gt; The catalog namespaces tables (&lt;code&gt;lake.orders&lt;/code&gt;) and is the authority every engine consults, so "same table" means "same catalog entry," not "same folder."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Option A — &lt;code&gt;SqlCatalog&lt;/code&gt; on SQLite (pure filesystem).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Zero infrastructure.&lt;/strong&gt; pyiceberg's &lt;code&gt;SqlCatalog&lt;/code&gt; stores the pointers in a local SQLite file and writes data + metadata to a &lt;code&gt;file://&lt;/code&gt; warehouse directory. No server, no Docker — perfect for a single-process test or a notebook.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trade-off.&lt;/strong&gt; SQLite is single-writer and not a network service, so it is ideal when one process owns the catalog but awkward when several tools must share it live.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Option B — a REST catalog (lakekeeper / iceberg-rest) on MinIO.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A shared metadata plane.&lt;/strong&gt; The Iceberg REST catalog is a standard HTTP API; run &lt;strong&gt;lakekeeper&lt;/strong&gt; (a Rust REST catalog) or the reference &lt;code&gt;iceberg-rest&lt;/code&gt; image in Docker and every engine points at one URL. This is how you get DuckDB, Trino, and pyiceberg reading and writing the &lt;em&gt;same&lt;/em&gt; table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MinIO as local S3.&lt;/strong&gt; MinIO is an S3-compatible object server you run in Docker (&lt;code&gt;minio server /data&lt;/code&gt;). Point the warehouse at &lt;code&gt;s3://warehouse&lt;/code&gt; with the MinIO endpoint, path-style access, and dummy &lt;code&gt;minioadmin&lt;/code&gt; keys — the exact S3 code path you use in production, exercised offline.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu82bfzzbtbs15v4t3xdm.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu82bfzzbtbs15v4t3xdm.jpeg" alt="Iconographic local Iceberg catalog diagram — a SQLite SqlCatalog and a REST catalog both pointing at an Iceberg table's metadata, with data files landing on either a file:// warehouse or a MinIO S3 bucket." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — create an Iceberg table with pyiceberg on a SQLite catalog
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The fastest way to get a real Iceberg table on a laptop is pyiceberg's &lt;code&gt;SqlCatalog&lt;/code&gt; backed by SQLite, writing to a &lt;code&gt;file://&lt;/code&gt; warehouse. You create a namespace, define a schema (or hand it a PyArrow table and let it infer), create the table, and append an Arrow batch. Everything lands on local disk as a proper Iceberg tree.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Using pyiceberg, create a &lt;code&gt;lake.orders&lt;/code&gt; Iceberg table on a local SQLite catalog and append two rows, with no cloud and no server.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; A two-row PyArrow table of orders.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pyarrow&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pa&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyiceberg.catalog.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SqlCatalog&lt;/span&gt;

&lt;span class="n"&gt;catalog&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SqlCatalog&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;          &lt;span class="c1"&gt;# pointers in SQLite, data in a local folder
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;local&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;uri&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sqlite:///warehouse/catalog.db&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;warehouse&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;file://warehouse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;catalog&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_namespace_if_not_exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lake&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pa&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;table&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;order_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ada&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;linus&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;   &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;42.50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;17.00&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="n"&gt;table&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;catalog&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_table_if_not_exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lake.orders&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;schema&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;schema&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;scan&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;to_arrow&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="n"&gt;num_rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# -&amp;gt; 2
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;SqlCatalog(...)&lt;/code&gt; opens (or creates) &lt;code&gt;warehouse/catalog.db&lt;/code&gt; as the pointer store and sets the warehouse root to the local &lt;code&gt;warehouse/&lt;/code&gt; folder. &lt;code&gt;create_namespace_if_not_exists("lake")&lt;/code&gt; registers the namespace. &lt;code&gt;create_table_if_not_exists("lake.orders", schema=rows.schema)&lt;/code&gt; writes the first &lt;code&gt;metadata.json&lt;/code&gt; under &lt;code&gt;warehouse/lake.db/orders/metadata/&lt;/code&gt; and records its path in SQLite. &lt;code&gt;table.append(rows)&lt;/code&gt; writes a Parquet data file, a new manifest and manifest list, a new &lt;code&gt;metadata.json&lt;/code&gt;, and atomically updates the SQLite pointer — one Iceberg snapshot. The final &lt;code&gt;scan()&lt;/code&gt; reads the current snapshot back as Arrow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;what pyiceberg created&lt;/th&gt;
&lt;th&gt;value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;catalog store&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;warehouse/catalog.db&lt;/code&gt; (SQLite)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;table root&lt;/td&gt;
&lt;td&gt;&lt;code&gt;warehouse/lake.db/orders/&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;data files&lt;/td&gt;
&lt;td&gt;one Parquet file under &lt;code&gt;.../data/&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;snapshots&lt;/td&gt;
&lt;td&gt;1 (the append)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;rows in current snapshot&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Use &lt;code&gt;SqlCatalog&lt;/code&gt; + SQLite when one process owns the table (unit tests, notebooks); switch to a REST catalog the moment a second engine needs to read or write the same table live.&lt;/p&gt;

&lt;h3&gt;
  
  
  Iceberg interview question on catalog choice
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; On a laptop, a pyiceberg job writes &lt;code&gt;lake.orders&lt;/code&gt; and a Trino instance must query it at the same time. A teammate suggests "just point both at the same warehouse folder — skip the catalog." Why is that wrong, and what do you set up instead?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using a shared REST catalog backed by MinIO
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;                      &lt;span class="c1"&gt;# docker-compose.yml — local control + storage plane&lt;/span&gt;
  &lt;span class="na"&gt;minio&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;                       &lt;span class="c1"&gt;# local S3&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;minio/minio&lt;/span&gt;
    &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;server /data --console-address ":9001"&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;MINIO_ROOT_USER&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;minioadmin&lt;/span&gt;
      &lt;span class="na"&gt;MINIO_ROOT_PASSWORD&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;minioadmin&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;9000:9000"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;9001:9001"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

  &lt;span class="na"&gt;lakekeeper&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;                  &lt;span class="c1"&gt;# Iceberg REST catalog&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;quay.io/lakekeeper/catalog:latest&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;LAKEKEEPER__PG_DATABASE_URL_READ_WRITE&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgres://catalog:catalog@pg/catalog&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8181:8181"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;depends_on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;pg&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;minio&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

  &lt;span class="na"&gt;pg&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgres:16&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;POSTGRES_USER&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;catalog&lt;/span&gt;
      &lt;span class="na"&gt;POSTGRES_PASSWORD&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;catalog&lt;/span&gt;
      &lt;span class="na"&gt;POSTGRES_DB&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;catalog&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyiceberg.catalog.rest&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;RestCatalog&lt;/span&gt;   &lt;span class="c1"&gt;# same URL for writer + reader
&lt;/span&gt;
&lt;span class="n"&gt;catalog&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RestCatalog&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;shared&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;uri&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:8181/catalog&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;warehouse&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;demo&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s3.endpoint&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:9000&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s3.access-key-id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;minioadmin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s3.secret-access-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;minioadmin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s3.path-style-access&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;approach&lt;/th&gt;
&lt;th&gt;who moves the metadata pointer&lt;/th&gt;
&lt;th&gt;concurrent reader sees&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;shared folder, no catalog&lt;/td&gt;
&lt;td&gt;nobody — no atomic commit&lt;/td&gt;
&lt;td&gt;half-written metadata, races&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;REST catalog (lakekeeper)&lt;/td&gt;
&lt;td&gt;the catalog, atomically&lt;/td&gt;
&lt;td&gt;old snapshot until commit lands&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;REST catalog + MinIO&lt;/td&gt;
&lt;td&gt;catalog for metadata, S3 for files&lt;/td&gt;
&lt;td&gt;consistent snapshot every read&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Iceberg's correctness depends on an &lt;strong&gt;atomic pointer swap&lt;/strong&gt; to the current &lt;code&gt;metadata.json&lt;/code&gt;; a bare folder has no component that performs that swap, so two processes writing or reading mid-commit can see a partial or inconsistent tree.&lt;/li&gt;
&lt;li&gt;A REST catalog (lakekeeper) owns that swap: the writer's &lt;code&gt;append&lt;/code&gt; is only visible once the catalog commits the new pointer, giving readers snapshot isolation instead of a race.&lt;/li&gt;
&lt;li&gt;Pointing the warehouse at MinIO (&lt;code&gt;s3.path-style-access=true&lt;/code&gt;, endpoint &lt;code&gt;:9000&lt;/code&gt;) exercises the real S3 code path locally, so the data files are addressed exactly as they would be in the cloud.&lt;/li&gt;
&lt;li&gt;Because every engine — pyiceberg writer, Trino reader, DuckDB &lt;code&gt;ATTACH&lt;/code&gt; — hits the &lt;strong&gt;same URL&lt;/strong&gt;, "the same table" is defined by one catalog entry, not by whoever happened to look at the folder.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;component&lt;/th&gt;
&lt;th&gt;local endpoint&lt;/th&gt;
&lt;th&gt;role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;lakekeeper&lt;/td&gt;
&lt;td&gt;&lt;code&gt;http://localhost:8181/catalog&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;atomic metadata pointer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MinIO&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;http://localhost:9000&lt;/code&gt; (&lt;code&gt;s3://warehouse&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;data + metadata files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;result&lt;/td&gt;
&lt;td&gt;one &lt;code&gt;lake.orders&lt;/code&gt; table&lt;/td&gt;
&lt;td&gt;shared, snapshot-isolated&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Catalog as source of truth&lt;/strong&gt;&lt;/strong&gt; — the catalog, not the folder, defines a table; only it can atomically move the current-metadata pointer that every engine reads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Atomic commit&lt;/strong&gt;&lt;/strong&gt; — the pointer swap is the transaction boundary; without it there is no isolation, and concurrent read/write over a shared folder corrupts what readers see.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;REST is engine-neutral&lt;/strong&gt;&lt;/strong&gt; — the Iceberg REST protocol is a standard, so DuckDB, Trino, and pyiceberg all speak it and converge on identical table state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;MinIO fidelity&lt;/strong&gt;&lt;/strong&gt; — running the real S3 API locally means the storage layer behaves like production (path-style, keys, endpoints), so tests catch S3-specific bugs offline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — a commit is O(1) metadata pointer update plus O(changed files) manifest writes, independent of table size, which is why commits stay cheap as data grows.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Warehouse&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — data-warehouse&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Lakehouse and warehouse-modelling problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-warehouse" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;ETL&lt;/span&gt;
&lt;span&gt;Topic — etl&lt;/span&gt;
&lt;strong&gt;Catalog-driven load and register problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Trino querying the same Iceberg tables
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Trino's &lt;code&gt;iceberg&lt;/code&gt; connector reads the catalogued table directly — same metadata, distributed SQL, time travel
&lt;/h3&gt;

&lt;p&gt;Once a table lives in a REST catalog, Trino queries it with its &lt;code&gt;iceberg&lt;/code&gt; connector — no copy, no import, just a catalog properties file that points Trino at the same REST endpoint and the same MinIO storage the writer used. This is what makes the local lakehouse a &lt;em&gt;lakehouse&lt;/em&gt; and not just "DuckDB on Parquet": full ANSI SQL, joins across tables, snapshot time travel, and DDL, all over the exact bytes pyiceberg wrote.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wiring Trino to the catalog.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Catalog properties file.&lt;/strong&gt; Drop &lt;code&gt;etc/catalog/iceberg.properties&lt;/code&gt; with &lt;code&gt;connector.name=iceberg&lt;/code&gt; and &lt;code&gt;iceberg.catalog.type=rest&lt;/code&gt;; Trino exposes it as the &lt;code&gt;iceberg&lt;/code&gt; catalog you reference in SQL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;REST endpoint.&lt;/strong&gt; &lt;code&gt;iceberg.rest-catalog.uri=http://lakekeeper:8181/catalog&lt;/code&gt; and &lt;code&gt;iceberg.rest-catalog.warehouse=demo&lt;/code&gt; tell Trino which catalog and warehouse to use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local S3 filesystem.&lt;/strong&gt; &lt;code&gt;fs.native-s3.enabled=true&lt;/code&gt;, &lt;code&gt;s3.endpoint=http://minio:9000&lt;/code&gt;, &lt;code&gt;s3.path-style-access=true&lt;/code&gt;, plus the &lt;code&gt;minioadmin&lt;/code&gt; keys, let Trino read the data files from MinIO exactly as pyiceberg wrote them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Querying and writing.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Plain SQL.&lt;/strong&gt; &lt;code&gt;SELECT * FROM iceberg.lake.orders&lt;/code&gt; reads the current snapshot; joins, window functions, and aggregates all work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CTAS and inserts.&lt;/strong&gt; &lt;code&gt;CREATE TABLE iceberg.lake.daily AS SELECT ...&lt;/code&gt; and &lt;code&gt;INSERT INTO&lt;/code&gt; create and grow Iceberg tables from Trino, committed through the same catalog.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metadata tables.&lt;/strong&gt; &lt;code&gt;SELECT * FROM iceberg.lake."orders$snapshots"&lt;/code&gt; and &lt;code&gt;"...$files"&lt;/code&gt; expose snapshots, manifests, and file lists for debugging.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Time travel and isolation.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;By timestamp.&lt;/strong&gt; &lt;code&gt;SELECT * FROM iceberg.lake.orders FOR TIMESTAMP AS OF TIMESTAMP '2026-02-01 00:00:00 UTC'&lt;/code&gt; reads the snapshot current at that instant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;By snapshot id.&lt;/strong&gt; &lt;code&gt;FOR VERSION AS OF &amp;lt;snapshot_id&amp;gt;&lt;/code&gt; pins an exact snapshot — invaluable for a deterministic test that must read a known state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Snapshot isolation.&lt;/strong&gt; A running &lt;code&gt;SELECT&lt;/code&gt; sees one snapshot even as a writer commits new ones, so readers never see a half-written table.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fytmhoz0i1n65zyybg02w.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fytmhoz0i1n65zyybg02w.jpeg" alt="Iconographic Trino-on-Iceberg diagram — a Trino coordinator with an iceberg connector reading the same REST-catalogued Iceberg table that pyiceberg wrote, including a time-travel snapshot selector." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — point Trino at the shared catalog and query it
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; With lakekeeper and MinIO running, Trino needs exactly one properties file to see the table pyiceberg created. After that, &lt;code&gt;SHOW TABLES&lt;/code&gt; and &lt;code&gt;SELECT&lt;/code&gt; work as if the table were native, because to Trino it &lt;em&gt;is&lt;/em&gt; native Iceberg.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Write the Trino &lt;code&gt;iceberg.properties&lt;/code&gt; for the local REST catalog + MinIO, then a query that returns per-customer revenue from &lt;code&gt;lake.orders&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; The &lt;code&gt;lake.orders&lt;/code&gt; table from Section 3 (columns &lt;code&gt;order_id&lt;/code&gt;, &lt;code&gt;customer&lt;/code&gt;, &lt;code&gt;amount&lt;/code&gt;), already committed to the catalog.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;connector.name&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;iceberg&lt;/span&gt;
&lt;span class="py"&gt;iceberg.catalog.type&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;rest&lt;/span&gt;
&lt;span class="py"&gt;iceberg.rest-catalog.uri&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;http://lakekeeper:8181/catalog&lt;/span&gt;
&lt;span class="py"&gt;iceberg.rest-catalog.warehouse&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;demo&lt;/span&gt;
&lt;span class="py"&gt;fs.native-s3.enabled&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;
&lt;span class="py"&gt;s3.endpoint&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;http://minio:9000&lt;/span&gt;
&lt;span class="py"&gt;s3.region&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;
&lt;span class="py"&gt;s3.path-style-access&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;
&lt;span class="py"&gt;s3.aws-access-key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;minioadmin&lt;/span&gt;
&lt;span class="py"&gt;s3.aws-secret-key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;minioadmin&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- trino&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;revenue&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;customer&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;revenue&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The properties file registers a Trino catalog named &lt;code&gt;iceberg&lt;/code&gt; whose &lt;code&gt;type=rest&lt;/code&gt; makes Trino fetch table metadata from lakekeeper at &lt;code&gt;:8181&lt;/code&gt;; the &lt;code&gt;fs.native-s3&lt;/code&gt; and &lt;code&gt;s3.*&lt;/code&gt; keys let Trino open the Parquet data files from MinIO with path-style addressing. In SQL, &lt;code&gt;iceberg.lake.orders&lt;/code&gt; resolves as &lt;code&gt;catalog.schema.table&lt;/code&gt; — Trino asks lakekeeper for the current snapshot, reads the manifests, and scans the data files. The &lt;code&gt;GROUP BY&lt;/code&gt; runs in Trino's engine and returns aggregated rows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;revenue&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;42.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;linus&lt;/td&gt;
&lt;td&gt;17.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Trino needs no data movement to read an Iceberg table — one properties file pointing at the shared catalog and storage is the entire integration, because the table format is the contract.&lt;/p&gt;

&lt;h3&gt;
  
  
  Trino interview question on cross-engine consistency and time travel
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; pyiceberg wrote &lt;code&gt;lake.orders&lt;/code&gt; at 10:00 (snapshot A) and appended more rows at 11:00 (snapshot B). A test must assert Trino reads exactly snapshot A's three rows, regardless of later writes. Write the query and explain why it is deterministic.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using Iceberg snapshot time travel in Trino
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- find the snapshots (newest first)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;snapshot_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;committed_at&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;"orders$snapshots"&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;committed_at&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- pin the exact snapshot the test expects (snapshot A = 10:00)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;rows_at_A&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="nb"&gt;TIMESTAMP&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;OF&lt;/span&gt; &lt;span class="nb"&gt;TIMESTAMP&lt;/span&gt; &lt;span class="s1"&gt;'2026-02-01 10:30:00 UTC'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- or pin by explicit id for full determinism&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;rows_at_A&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="k"&gt;VERSION&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;OF&lt;/span&gt; &lt;span class="mi"&gt;4823174097&lt;/span&gt; &lt;span class="cm"&gt;/* snapshot A id */&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;clock&lt;/th&gt;
&lt;th&gt;snapshot current&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;FOR TIMESTAMP AS OF 10:30&lt;/code&gt; resolves to&lt;/th&gt;
&lt;th&gt;rows seen&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10:00&lt;/td&gt;
&lt;td&gt;A committed&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11:00&lt;/td&gt;
&lt;td&gt;B committed&lt;/td&gt;
&lt;td&gt;still A (10:30 &amp;lt; 11:00)&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;now&lt;/td&gt;
&lt;td&gt;B is latest&lt;/td&gt;
&lt;td&gt;A, because 10:30 predates B&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The &lt;code&gt;orders$snapshots&lt;/code&gt; metadata table lists every commit with its &lt;code&gt;snapshot_id&lt;/code&gt; and &lt;code&gt;committed_at&lt;/code&gt;, so the test can discover or hard-code the snapshot it wants.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;FOR TIMESTAMP AS OF TIMESTAMP '... 10:30 ...'&lt;/code&gt; tells Trino to resolve the snapshot that was current at 10:30 — that is snapshot A, because B did not commit until 11:00.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;FOR VERSION AS OF &amp;lt;snapshot_id&amp;gt;&lt;/code&gt; skips the timestamp resolution entirely and reads an exact snapshot, which is the most deterministic form for CI.&lt;/li&gt;
&lt;li&gt;Because Iceberg snapshots are immutable, snapshot A's three rows can never change no matter how many later appends land, so the assertion holds on every run.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;rows_at_A&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Immutable snapshots&lt;/strong&gt;&lt;/strong&gt; — every commit creates a new snapshot and never mutates an old one, so a pinned snapshot is a frozen, reproducible view of the table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Time travel&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;FOR TIMESTAMP AS OF&lt;/code&gt; / &lt;code&gt;FOR VERSION AS OF&lt;/code&gt; let a query address a past state by clock or id, which is exactly what a deterministic test needs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Shared catalog&lt;/strong&gt;&lt;/strong&gt; — Trino resolves the snapshot through the same catalog pyiceberg wrote to, so "snapshot A" means the same thing to both engines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Snapshot isolation&lt;/strong&gt;&lt;/strong&gt; — a reader is bound to one snapshot for the whole query, so concurrent writes never leak partial rows into the result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — time travel is O(1) to resolve the pointer plus O(scanned files) for that snapshot; it reads an old state at no extra bookkeeping cost because the files already exist.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Optimization&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — optimization&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Query-planning and scan-optimization problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/optimization" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Partitioning&lt;/span&gt;
&lt;span&gt;Topic — partitioning&lt;/span&gt;
&lt;strong&gt;Partition-pruning and table-layout problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/partitioning" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Partitioning and compaction on your laptop
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Hidden partitioning plus &lt;code&gt;EXECUTE optimize&lt;/code&gt; fixes the small-files problem locally — the same maintenance you run in production
&lt;/h3&gt;

&lt;p&gt;The last piece that makes a local lakehouse behave like production is table maintenance: partitioning data so scans prune, and compacting the many small files that iterative local writes produce. Iceberg does both with declarative DDL and table procedures, and you run the &lt;em&gt;identical&lt;/em&gt; operations locally that you would run against a cloud warehouse — so a test can prove your partition spec and compaction actually help.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hidden partitioning.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Declare a spec, not a column.&lt;/strong&gt; &lt;code&gt;WITH (partitioning = ARRAY['day(order_ts)'])&lt;/code&gt; partitions by a &lt;em&gt;derived&lt;/em&gt; value; you query on &lt;code&gt;order_ts&lt;/code&gt; and Iceberg prunes without you materialising a &lt;code&gt;day&lt;/code&gt; column. That is "hidden" partitioning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transforms.&lt;/strong&gt; &lt;code&gt;day(ts)&lt;/code&gt;, &lt;code&gt;month(ts)&lt;/code&gt;, &lt;code&gt;hour(ts)&lt;/code&gt;, &lt;code&gt;bucket(N, id)&lt;/code&gt;, and &lt;code&gt;truncate(N, col)&lt;/code&gt; are the built-in partition transforms; &lt;code&gt;bucket(16, user_id)&lt;/code&gt; spreads a high-cardinality key into 16 even buckets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evolvable.&lt;/strong&gt; You can change the partition spec later without rewriting history — old data keeps its old spec, new data uses the new one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The small-files problem.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Why it appears.&lt;/strong&gt; Each write commits at least one data file per partition; a test loop or a streaming append produces thousands of tiny Parquet files, and every query pays per-file open and metadata overhead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The symptom.&lt;/strong&gt; Scans that should be fast crawl, because the engine spends its time opening files and reading footers rather than scanning data.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Compaction and housekeeping (Trino procedures).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Compact.&lt;/strong&gt; &lt;code&gt;ALTER TABLE iceberg.lake.orders EXECUTE optimize&lt;/code&gt; rewrites small files into fewer, right-sized ones; &lt;code&gt;EXECUTE optimize(file_size_threshold =&amp;gt; '128MB')&lt;/code&gt; only compacts files below the threshold.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expire snapshots.&lt;/strong&gt; &lt;code&gt;ALTER TABLE ... EXECUTE expire_snapshots(retention_threshold =&amp;gt; '7d')&lt;/code&gt; drops old snapshots and the data files only they referenced, reclaiming space.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remove orphans.&lt;/strong&gt; &lt;code&gt;ALTER TABLE ... EXECUTE remove_orphan_files(retention_threshold =&amp;gt; '7d')&lt;/code&gt; deletes files no live snapshot references — the leftovers of failed writes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why it matters even locally.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CI stays fast and deterministic.&lt;/strong&gt; Compacting fixtures keeps test scans quick; expiring snapshots keeps the local warehouse from growing unbounded across CI runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You test the maintenance itself.&lt;/strong&gt; Running &lt;code&gt;optimize&lt;/code&gt; locally lets you assert file counts dropped and query time fell — you are validating the production maintenance job, not guessing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffmb81mrxr7ocp4o6xvkc.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffmb81mrxr7ocp4o6xvkc.jpeg" alt="Iconographic Iceberg partitioning and compaction diagram — hidden partitioning by day and bucket, many small files rewritten into a few right-sized files by EXECUTE optimize, and expired snapshots pruned." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — a partitioned table, then compact it
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The clearest local demonstration is to create a partitioned Iceberg table in Trino, write to it a few times so it accumulates small files, then run &lt;code&gt;EXECUTE optimize&lt;/code&gt; and watch the file count collapse while the row count stays identical. Nothing leaves your laptop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Create &lt;code&gt;lake.events&lt;/code&gt; partitioned by &lt;code&gt;day(event_ts)&lt;/code&gt; and &lt;code&gt;bucket(8, user_id)&lt;/code&gt;, then compact it after several small appends. Show the file count before and after.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; Three appends of ~1,000 rows each, landing many small files across partitions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- trino: create with hidden partitioning&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;user_id&lt;/span&gt;   &lt;span class="nb"&gt;BIGINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;event_ts&lt;/span&gt;  &lt;span class="nb"&gt;TIMESTAMP&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;amount&lt;/span&gt;    &lt;span class="nb"&gt;DOUBLE&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;partitioning&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ARRAY&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'day(event_ts)'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'bucket(8, user_id)'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;format&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'PARQUET'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- ... three separate INSERT INTO iceberg.lake.events SELECT ... appends ...&lt;/span&gt;

&lt;span class="c1"&gt;-- inspect small files, then compact&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;data_files&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;"events$files"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;EXECUTE&lt;/span&gt; &lt;span class="n"&gt;optimize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_size_threshold&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s1"&gt;'128MB'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;data_files&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;"events$files"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The &lt;code&gt;partitioning = ARRAY['day(event_ts)', 'bucket(8, user_id)']&lt;/code&gt; clause records a partition spec of a date-day transform plus an 8-way hash bucket, so data is physically grouped by day and user-bucket. Each &lt;code&gt;INSERT&lt;/code&gt; commits at least one file per touched partition, so three appends across many partitions leave lots of small files, which &lt;code&gt;events$files&lt;/code&gt; counts. &lt;code&gt;EXECUTE optimize(file_size_threshold =&amp;gt; '128MB')&lt;/code&gt; rewrites all files below 128 MB within each partition into fewer, larger files in a new snapshot, without changing any row. The second count shows the collapse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;stage&lt;/th&gt;
&lt;th&gt;data files&lt;/th&gt;
&lt;th&gt;rows&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;after 3 appends&lt;/td&gt;
&lt;td&gt;42&lt;/td&gt;
&lt;td&gt;3000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;after &lt;code&gt;EXECUTE optimize&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;3000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Partition by what you filter on, bucket high-cardinality join keys, and schedule &lt;code&gt;EXECUTE optimize&lt;/code&gt; after bursty writes — locally you can &lt;em&gt;prove&lt;/em&gt; the file count dropped, so you ship the maintenance job with confidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Iceberg interview question on local compaction and snapshot hygiene
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A local CI job writes a fixture in a loop and leaves 10,000 tiny files plus dozens of stale snapshots; Trino scans have become slow and the warehouse folder keeps growing between runs. Fix both the scan speed and the growth, and justify why this matters in CI, not just production.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using EXECUTE optimize plus expire_snapshots
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- 1) compact small files into right-sized ones (new snapshot)&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt;
&lt;span class="k"&gt;EXECUTE&lt;/span&gt; &lt;span class="n"&gt;optimize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_size_threshold&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s1"&gt;'128MB'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- 2) drop snapshots older than the retention window and their&lt;/span&gt;
&lt;span class="c1"&gt;--    now-unreferenced data files&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt;
&lt;span class="k"&gt;EXECUTE&lt;/span&gt; &lt;span class="n"&gt;expire_snapshots&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;retention_threshold&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s1"&gt;'7d'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- 3) sweep files no live snapshot references (failed-write leftovers)&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt;
&lt;span class="k"&gt;EXECUTE&lt;/span&gt; &lt;span class="n"&gt;remove_orphan_files&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;retention_threshold&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s1"&gt;'7d'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;procedure&lt;/th&gt;
&lt;th&gt;effect on files&lt;/th&gt;
&lt;th&gt;effect on snapshots&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;optimize&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10,000 small → ~40 sized&lt;/td&gt;
&lt;td&gt;+1 (compaction snapshot)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;&lt;code&gt;expire_snapshots&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;deletes files only old snapshots used&lt;/td&gt;
&lt;td&gt;drops stale snapshots&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;remove_orphan_files&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;deletes unreferenced leftovers&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;EXECUTE optimize&lt;/code&gt; rewrites the 10,000 tiny files into a handful of ~128 MB files, so every subsequent scan opens dozens of files instead of ten thousand — the scan-speed fix.&lt;/li&gt;
&lt;li&gt;Compaction itself adds a new snapshot and leaves the old small files still referenced by earlier snapshots, so space does not drop yet; that is expected.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;EXECUTE expire_snapshots&lt;/code&gt; removes snapshots past the retention window and physically deletes the data files that only those expired snapshots referenced — this is where the folder finally shrinks.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;EXECUTE remove_orphan_files&lt;/code&gt; deletes files on storage that no live snapshot points to at all (typically from interrupted writes), reclaiming the last of the wasted space.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;metric&lt;/th&gt;
&lt;th&gt;before&lt;/th&gt;
&lt;th&gt;after maintenance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;data files&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;~40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;live snapshots&lt;/td&gt;
&lt;td&gt;60+&lt;/td&gt;
&lt;td&gt;a few (within 7d)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;warehouse size&lt;/td&gt;
&lt;td&gt;bloated&lt;/td&gt;
&lt;td&gt;lean&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trino scan time&lt;/td&gt;
&lt;td&gt;slow&lt;/td&gt;
&lt;td&gt;fast&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Compaction&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;optimize&lt;/code&gt; trades a one-time rewrite for cheap steady-state scans by replacing many small files with few right-sized ones, eliminating per-file open overhead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Snapshot retention&lt;/strong&gt;&lt;/strong&gt; — old snapshots pin their data files alive; &lt;code&gt;expire_snapshots&lt;/code&gt; is what actually frees space, because compaction alone only adds a snapshot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Orphan cleanup&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;remove_orphan_files&lt;/code&gt; reclaims files no snapshot references, which accumulate from failed or interrupted local writes during a test loop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;CI relevance&lt;/strong&gt;&lt;/strong&gt; — a fixture that bloats every run makes CI slower and flakier over time; running the real maintenance locally keeps runs fast &lt;em&gt;and&lt;/em&gt; validates the production job.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;optimize&lt;/code&gt; is O(rewritten bytes) once; the payoff is O(files) reduced per query forever after, and the housekeeping procedures are O(expired files) to delete.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Partitioning&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — partitioning&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Partition-spec and file-sizing problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/partitioning" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Optimization&lt;/span&gt;
&lt;span&gt;Topic — optimization&lt;/span&gt;
&lt;strong&gt;Compaction and scan-cost optimization problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/optimization" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — local lakehouse recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;DuckDB — write and read Parquet.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;COPY&lt;/span&gt; &lt;span class="n"&gt;tbl&lt;/span&gt; &lt;span class="k"&gt;TO&lt;/span&gt; &lt;span class="s1"&gt;'lake/orders'&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;FORMAT&lt;/span&gt; &lt;span class="n"&gt;PARQUET&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;PARTITION_BY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event_date&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;read_parquet&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'lake/orders/**/*.parquet'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hive_partitioning&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;event_date&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="s1"&gt;'2026-02-11'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;DuckDB — read an Iceberg table.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;INSTALL&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;LOAD&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;iceberg_scan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'warehouse/lake.db/orders'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;allow_moved_paths&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="c1"&gt;-- or attach the shared REST catalog:&lt;/span&gt;
&lt;span class="n"&gt;ATTACH&lt;/span&gt; &lt;span class="s1"&gt;''&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;ice&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;TYPE&lt;/span&gt; &lt;span class="n"&gt;ICEBERG&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ENDPOINT&lt;/span&gt; &lt;span class="s1"&gt;'http://localhost:8181/catalog'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;ice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;pyiceberg — create + append on a SQLite catalog.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyiceberg.catalog.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SqlCatalog&lt;/span&gt;
&lt;span class="n"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SqlCatalog&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;local&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;uri&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sqlite:///warehouse/catalog.db&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;warehouse&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;file://warehouse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;cat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_namespace_if_not_exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lake&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_table_if_not_exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lake.orders&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;schema&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;arrow_table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;schema&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;arrow_table&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Trino — Iceberg REST catalog on MinIO.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;connector.name&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;iceberg&lt;/span&gt;
&lt;span class="py"&gt;iceberg.catalog.type&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;rest&lt;/span&gt;
&lt;span class="py"&gt;iceberg.rest-catalog.uri&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;http://lakekeeper:8181/catalog&lt;/span&gt;
&lt;span class="py"&gt;iceberg.rest-catalog.warehouse&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;demo&lt;/span&gt;
&lt;span class="py"&gt;fs.native-s3.enabled&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;
&lt;span class="py"&gt;s3.endpoint&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;http://minio:9000&lt;/span&gt;
&lt;span class="py"&gt;s3.path-style-access&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Trino — partitioned table + time travel.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_ts&lt;/span&gt; &lt;span class="nb"&gt;TIMESTAMP&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="nb"&gt;DOUBLE&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;partitioning&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ARRAY&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'day(event_ts)'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'bucket(8, user_id)'&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="nb"&gt;TIMESTAMP&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;OF&lt;/span&gt; &lt;span class="nb"&gt;TIMESTAMP&lt;/span&gt; &lt;span class="s1"&gt;'2026-02-01 10:30:00 UTC'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Trino — compaction and housekeeping.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;EXECUTE&lt;/span&gt; &lt;span class="n"&gt;optimize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_size_threshold&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s1"&gt;'128MB'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;EXECUTE&lt;/span&gt; &lt;span class="n"&gt;expire_snapshots&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;retention_threshold&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s1"&gt;'7d'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;iceberg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;EXECUTE&lt;/span&gt; &lt;span class="n"&gt;remove_orphan_files&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;retention_threshold&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s1"&gt;'7d'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Docker — the storage + catalog plane.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-p&lt;/span&gt; 9000:9000 &lt;span class="nt"&gt;-p&lt;/span&gt; 9001:9001 minio/minio server /data &lt;span class="nt"&gt;--console-address&lt;/span&gt; &lt;span class="s2"&gt;":9001"&lt;/span&gt;
docker run &lt;span class="nt"&gt;-p&lt;/span&gt; 8181:8181 quay.io/lakekeeper/catalog:latest   &lt;span class="c"&gt;# Iceberg REST catalog&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Engine roles at a glance.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Role in the local lakehouse&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DuckDB&lt;/td&gt;
&lt;td&gt;in-process reader + Parquet producer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pyiceberg&lt;/td&gt;
&lt;td&gt;Python writer, catalog + table author&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lakekeeper / iceberg-rest&lt;/td&gt;
&lt;td&gt;shared Iceberg REST catalog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MinIO&lt;/td&gt;
&lt;td&gt;local S3-compatible object storage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trino&lt;/td&gt;
&lt;td&gt;distributed SQL, DDL, compaction, time travel&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is a local lakehouse?
&lt;/h3&gt;

&lt;p&gt;A local lakehouse is a full lakehouse stack — an open table format (Apache Iceberg), object or file storage, and query engines — running entirely on your own machine with no cloud account. You land Parquet files on disk or in a local MinIO bucket, register them as Iceberg tables in a catalog you host (a SQLite &lt;code&gt;SqlCatalog&lt;/code&gt; or a REST catalog like lakekeeper), and query them from DuckDB or Trino. Because the table format is identical to production Iceberg, local dev and CI exercise the real semantics, not a stand-in.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can DuckDB write Iceberg tables?
&lt;/h3&gt;

&lt;p&gt;DuckDB's &lt;code&gt;iceberg&lt;/code&gt; extension has mature &lt;em&gt;read&lt;/em&gt; support — &lt;code&gt;iceberg_scan(...)&lt;/code&gt; reads a table by its metadata path, and recent versions can &lt;code&gt;ATTACH&lt;/code&gt; an Iceberg REST catalog and query its tables. Iceberg &lt;em&gt;write&lt;/em&gt; support from DuckDB is newer and still maturing, so most local setups author and mutate Iceberg tables with pyiceberg or Trino and use DuckDB as a fast reader and a Parquet producer. DuckDB writing plain Parquet, by contrast, is fully supported and is how you feed data into the lakehouse.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I need MinIO, or is the local filesystem enough?
&lt;/h3&gt;

&lt;p&gt;For a single process — a unit test or a notebook — a &lt;code&gt;file://&lt;/code&gt; warehouse with a SQLite &lt;code&gt;SqlCatalog&lt;/code&gt; is enough and needs no Docker at all. Add MinIO when you want to exercise the real S3 code path (path-style addressing, endpoints, credentials) or when multiple engines must share storage. MinIO is S3-compatible, so the same &lt;code&gt;s3://&lt;/code&gt; configuration you test locally works unchanged against real S3 in production.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do DuckDB, Trino, and pyiceberg share the same Iceberg table?
&lt;/h3&gt;

&lt;p&gt;They share a &lt;strong&gt;catalog&lt;/strong&gt;. The catalog holds the current metadata pointer for each table, and every engine that points at the same catalog URL sees the same table state. Run an Iceberg REST catalog (lakekeeper or the reference &lt;code&gt;iceberg-rest&lt;/code&gt;) and configure pyiceberg's &lt;code&gt;RestCatalog&lt;/code&gt;, Trino's &lt;code&gt;iceberg.catalog.type=rest&lt;/code&gt;, and DuckDB's &lt;code&gt;ATTACH ... (TYPE ICEBERG, ENDPOINT ...)&lt;/code&gt; to the one URL. Sharing a folder without a catalog does not work, because only the catalog performs the atomic commit that gives snapshot isolation.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I compact small files in a local Iceberg table?
&lt;/h3&gt;

&lt;p&gt;Run &lt;code&gt;ALTER TABLE iceberg.lake.events EXECUTE optimize(file_size_threshold =&amp;gt; '128MB')&lt;/code&gt; in Trino to rewrite small files into fewer right-sized ones. Follow it with &lt;code&gt;EXECUTE expire_snapshots(retention_threshold =&amp;gt; '7d')&lt;/code&gt; to drop old snapshots and reclaim the space their data files held, and &lt;code&gt;EXECUTE remove_orphan_files(...)&lt;/code&gt; to sweep leftovers from failed writes. Compaction alone does not shrink storage — you must expire snapshots for the old files to be deleted.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why run a lakehouse locally instead of a dev cloud account?
&lt;/h3&gt;

&lt;p&gt;Speed, determinism, and cost. A local loop of write-query-verify is seconds, not the minutes a shared warehouse round-trip costs, and it needs no credentials or network. Iceberg snapshot isolation plus a fixed local dataset make tests reproducible, so CI stops being flaky from shared-state collisions. And because CI runners spin the lakehouse up inside the job, there is no cloud bill, no secret management, and no egress — while still testing the exact table format, partitioning, and compaction you run in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every local-lakehouse idea above, from DuckDB Parquet pushdown to the shared Iceberg REST catalog, Trino time travel, and `EXECUTE optimize` compaction, maps to a hands-on practice room where you build the load against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you make this scan prune and this table stay compact?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/etl" rel="noopener noreferrer"&gt;Practice ETL problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/optimization" rel="noopener noreferrer"&gt;Optimization drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Data Residency &amp; Sovereignty: Multi-Region Architectures for Compliant Data Platforms</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Sun, 27 Sep 2026 17:37:05 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/data-residency-sovereignty-multi-region-architectures-for-compliant-data-platforms-289b</link>
      <guid>https://dev.to/gowthampotureddi/data-residency-sovereignty-multi-region-architectures-for-compliant-data-platforms-289b</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;data residency&lt;/code&gt;&lt;/strong&gt; is a claim about geography — it says the physical bytes of a given dataset are stored and processed inside a named region, and never leave it. That sounds like a one-line configuration flag, and for a single bucket it almost is. But a real data platform is a mesh of storage layers, warehouses, replication jobs, caches, backups, key stores, and query engines, and residency is only satisfied when &lt;em&gt;every&lt;/em&gt; one of those components keeps the row inside the boundary. The moment a nightly cross-region replica, a global materialized view, or a helpfully-replicated encryption key carries a byte across the line, the guarantee is gone — usually silently, and usually discovered in an audit rather than a test.&lt;/p&gt;

&lt;p&gt;Residency is also routinely confused with two neighbours it is not. It is not &lt;strong&gt;&lt;code&gt;data sovereignty&lt;/code&gt;&lt;/strong&gt;, which is the harder question of &lt;em&gt;whose laws&lt;/em&gt; can compel access to the data regardless of where it sits — a US-headquartered provider can be reachable under the US CLOUD Act even for bytes stored in Frankfurt. And it is not &lt;strong&gt;disaster recovery&lt;/strong&gt;, which is about keeping &lt;em&gt;copies&lt;/em&gt; for availability; DR wants your data in more than one place, residency wants it in exactly one jurisdiction, and the two goals pull in opposite directions. This guide walks the four design problems an interviewer will actually probe — region-pinned storage with geo-partitioning, per-region warehouse deployments with hard replication boundaries, encryption with in-region key residency, and cross-region aggregate-only egress — and pairs each with a Solution-Tail answer: code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw2lroj4jjfzmjloszvnn.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw2lroj4jjfzmjloszvnn.jpeg" alt="PipeCode blog header for data residency and sovereignty — bold white headline 'Residency &amp;amp; Sovereignty' with subtitle 'Multi-Region Architectures for Compliant Data Platforms' and a stylised world-regions scene with region-pinned data vaults and a blocked cross-region arrow on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/partitioning" rel="noopener noreferrer"&gt;partitioning practice library →&lt;/a&gt;, rehearse the region-shard decisions on the &lt;a href="https://pipecode.ai/explore/practice/topic/sharding" rel="noopener noreferrer"&gt;sharding practice set →&lt;/a&gt;, and harden who-can-read-what on the &lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;access-control practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why residency and sovereignty are different problems&lt;/li&gt;
&lt;li&gt;Region-pinned storage &amp;amp; geo-partitioning&lt;/li&gt;
&lt;li&gt;Regional warehouses &amp;amp; replication boundaries&lt;/li&gt;
&lt;li&gt;Encryption, key residency &amp;amp; BYOK&lt;/li&gt;
&lt;li&gt;Cross-region aggregate-only egress&lt;/li&gt;
&lt;li&gt;Cheat sheet — residency recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why residency and sovereignty are different problems
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Residency is about geography, sovereignty is about jurisdiction — conflate them and you build the wrong control
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;residency answers "where is the data physically stored and processed?" while sovereignty answers "whose legal authority can reach it?" — and a design can satisfy one while violating the other&lt;/strong&gt;. You can host EU customer data in an EU region (residency satisfied) on a provider whose parent company is legally compellable under a foreign statute (sovereignty not satisfied). Interviewers who work in regulated industries open with this distinction on purpose, because a candidate who treats the two as synonyms will pick a control that looks compliant on an architecture diagram but fails the actual legal test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The two questions, kept apart.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Residency — where the bytes live.&lt;/strong&gt; A physical-location property of every copy of the data: primary storage, replicas, backups, temp files, query spill, log lines that echo payloads, and CDN caches. Residency is satisfied only when the &lt;em&gt;union&lt;/em&gt; of all those locations stays inside the allowed region set.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sovereignty — whose laws govern.&lt;/strong&gt; A jurisdictional property: which government can lawfully compel disclosure or seizure. This depends on where the operator is incorporated, where its staff and support sit, and which treaties apply — not only where the disk is. The US CLOUD Act and various national-security statutes are the usual examples of extraterritorial reach.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why the gap matters.&lt;/strong&gt; The Schrems II ruling turned on exactly this gap: storing EU data in the EU was not sufficient if a non-EU authority could still compel it. Sovereignty controls (in-region operator, customer-held keys, confidential computing) exist precisely to close a gap that residency alone leaves open.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Distinct from disaster recovery — the copies problem.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DR wants redundancy across failure domains.&lt;/strong&gt; The instinct of every availability design is "replicate to a second region so a regional outage does not take us down." That instinct directly manufactures a residency violation if the second region is in a different jurisdiction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Residency wants a bounded footprint.&lt;/strong&gt; The reconciliation is to keep DR &lt;em&gt;inside&lt;/em&gt; the residency boundary — a second availability zone or a second in-jurisdiction region — never a convenient far-away region. "Back it up to us-east-1 for safety" is the classic way an EU-residency platform quietly breaks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backups and snapshots count.&lt;/strong&gt; A backup is a copy; a copy has a location; that location is in scope. Cross-region backup replication is the single most common silent residency breach because it is configured once, for good reasons, and then forgotten.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The three legal forces that drive the architecture.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Data-localization laws.&lt;/strong&gt; Some jurisdictions require certain data to be stored (and sometimes processed) &lt;em&gt;only&lt;/em&gt; inside the country — for example national rules covering personal, health, financial, or government data. These force region-pinning and per-region deployments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-border transfer rules.&lt;/strong&gt; Frameworks like the GDPR permit transfers only under specific mechanisms (adequacy decisions, standard contractual clauses, supplementary measures). These force you to prove what leaves a region and under what basis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extraterritorial reach.&lt;/strong&gt; Statutes that let a government compel a provider regardless of storage location. These force sovereignty controls that residency cannot provide — customer-controlled keys and, at the extreme, sovereign-cloud operators.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you separate &lt;strong&gt;"where it is stored"&lt;/strong&gt; from &lt;strong&gt;"who can compel it"&lt;/strong&gt; in the first two sentences? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you flag that &lt;strong&gt;DR replication is the usual way residency breaks&lt;/strong&gt;, unprompted? — required framing.&lt;/li&gt;
&lt;li&gt;Do you name &lt;strong&gt;backups, temp files, and logs&lt;/strong&gt; as in-scope copies, not just the primary table? — the detail that separates a real design from a diagram.&lt;/li&gt;
&lt;li&gt;Do you reach for &lt;strong&gt;customer-held keys&lt;/strong&gt; when the interviewer adds "but the operator is foreign"? — the sovereignty move.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — the same customer, two very different guarantees
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The cleanest way to feel the distinction is to hold the storage location fixed and vary only the operator, then hold the operator fixed and vary only the storage. Each move flips exactly one of the two properties, which is the whole point: residency and sovereignty are independent axes, and a compliant design has to satisfy both simultaneously rather than assuming one implies the other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; For an EU customer whose personal data must stay in the EU and must not be reachable by a non-EU authority, classify four deployment options as residency-OK / sovereignty-OK.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;option&lt;/th&gt;
&lt;th&gt;storage region&lt;/th&gt;
&lt;th&gt;operator jurisdiction&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;EU (Frankfurt)&lt;/td&gt;
&lt;td&gt;non-EU parent company&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;US (Virginia)&lt;/td&gt;
&lt;td&gt;non-EU parent company&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;EU (Frankfurt)&lt;/td&gt;
&lt;td&gt;EU sovereign operator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D&lt;/td&gt;
&lt;td&gt;EU (Frankfurt)&lt;/td&gt;
&lt;td&gt;EU operator, customer-held keys&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;residency_ok  = (storage_region in ALLOWED_REGIONS)
sovereignty_ok = (operator_in_jurisdiction) OR (customer_controls_keys)

A: residency_ok=True,  sovereignty_ok=False   # in-region, but compellable abroad
B: residency_ok=False, sovereignty_ok=False   # wrong region entirely
C: residency_ok=True,  sovereignty_ok=True    # in-region + in-jurisdiction operator
D: residency_ok=True,  sovereignty_ok=True    # in-region + crypto control of access
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Option A stores in the EU so residency passes, but the foreign parent can be compelled, so sovereignty fails — this is the Schrems II shape exactly. Option B fails both because the bytes are in the wrong region. Option C satisfies both by using an operator that is itself inside the jurisdiction. Option D keeps a commercial operator but moves the &lt;em&gt;access&lt;/em&gt; control to the customer via customer-held encryption keys, so even a lawful demand to the operator yields ciphertext — sovereignty satisfied by cryptography rather than by corporate structure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;option&lt;/th&gt;
&lt;th&gt;residency&lt;/th&gt;
&lt;th&gt;sovereignty&lt;/th&gt;
&lt;th&gt;verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;OK&lt;/td&gt;
&lt;td&gt;FAIL&lt;/td&gt;
&lt;td&gt;non-compliant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;FAIL&lt;/td&gt;
&lt;td&gt;FAIL&lt;/td&gt;
&lt;td&gt;non-compliant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;OK&lt;/td&gt;
&lt;td&gt;OK&lt;/td&gt;
&lt;td&gt;compliant (sovereign operator)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D&lt;/td&gt;
&lt;td&gt;OK&lt;/td&gt;
&lt;td&gt;OK&lt;/td&gt;
&lt;td&gt;compliant (customer keys)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Fix the storage region first to satisfy residency, then ask "who could still be compelled?" to satisfy sovereignty — if the answer is a party you do not control, you need in-region operation or customer-held keys, not a second region.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Region-pinned storage &amp;amp; geo-partitioning
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Pin the bucket to a region and partition the fact table on a region key — location becomes a column, not a hope
&lt;/h3&gt;

&lt;p&gt;The feature that makes residency &lt;em&gt;enforceable&lt;/em&gt; rather than &lt;em&gt;aspirational&lt;/em&gt; is that in modern storage the region is an immutable property of the container and the data model carries the region as a first-class key. You pin the S3 bucket / GCS bucket / warehouse dataset to a region at creation, you disable any replication that would copy it elsewhere, and you make &lt;code&gt;region&lt;/code&gt; a partition key on every table that holds regulated rows so a row's residency is visible in the schema and enforceable in every query.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pinning the storage container.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Region is set at creation and often immutable.&lt;/strong&gt; An S3 bucket lives in one region for life; a BigQuery dataset's &lt;code&gt;location&lt;/code&gt; (&lt;code&gt;EU&lt;/code&gt;, &lt;code&gt;US&lt;/code&gt;, &lt;code&gt;europe-west3&lt;/code&gt;) is fixed when created and cannot be changed — you recreate to move. Treat the region as part of the resource's identity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Block the replication features by default.&lt;/strong&gt; S3 Cross-Region Replication, cross-region snapshot copy, and global-table features are opt-in conveniences that break residency. For a residency-bound bucket, the correct posture is an explicit deny (a bucket policy / SCP that forbids replication to non-allowed regions), not merely "we did not turn it on."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Watch the implicit copies.&lt;/strong&gt; Query engines spill to temp storage, CDNs cache objects at edges worldwide, and logs can echo payloads. Each of those has a region; each must be constrained to the boundary or scrubbed of regulated content.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Geo-partitioning the data model.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Make &lt;code&gt;region&lt;/code&gt; (or &lt;code&gt;country&lt;/code&gt;) a partition key.&lt;/strong&gt; &lt;code&gt;PARTITION BY region&lt;/code&gt; (or a Hive-style &lt;code&gt;region=EU/&lt;/code&gt; path layout on object storage) physically groups each region's rows into their own files / partitions, so residency is a property you can see, scan, and enforce per partition.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Partition pruning routes the query.&lt;/strong&gt; A query filtered &lt;code&gt;WHERE region = 'EU'&lt;/code&gt; reads only the EU partitions — the US partitions are pruned and never touched. This is both a performance win and a residency control: an EU-scoped job provably cannot read US data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bucketing within a region for skew.&lt;/strong&gt; If one region is huge, add a secondary &lt;code&gt;bucket&lt;/code&gt;/&lt;code&gt;shard&lt;/code&gt; on a high-cardinality key (customer_id) &lt;em&gt;inside&lt;/em&gt; the region so files stay evenly sized without ever mixing regions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Row-domiciled databases (the OLTP side).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CockroachDB &lt;code&gt;REGIONAL BY ROW&lt;/code&gt;.&lt;/strong&gt; Each row carries a hidden &lt;code&gt;crdb_region&lt;/code&gt; column; the database physically homes the row's replicas in that region's nodes, so a single logical table transparently keeps EU rows in EU nodes and US rows in US nodes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Citus / Yugabyte tablespaces.&lt;/strong&gt; Distribute a Postgres table by a region/tenant key and pin each shard's tablespace to region-local storage — the same "location is a key" idea in a Postgres-compatible engine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The invariant.&lt;/strong&gt; Whether OLTP or lakehouse, the design rule is identical: the region is part of the primary/partition key, and the storage for each key value is physically in that region.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjmjpc6hrjvv7et2etsxr.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjmjpc6hrjvv7et2etsxr.jpeg" alt="Iconographic diagram of region-pinned storage and geo-partitioning — a fact table partitioned on a region key into EU and US partition boxes, each pinned to a bucket whose region is locked, with cross-region replication blocked and a query pruned to a single region's partitions." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — geo-partitioning a fact table and proving the prune
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday pattern is a fact table partitioned on &lt;code&gt;region&lt;/code&gt; where each partition maps to region-local storage. The proof that residency holds is that a region-scoped query only ever scans that region's partition — you can read the guarantee straight out of the query plan, which is exactly what an auditor (or interviewer) wants to see.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Partition an &lt;code&gt;events&lt;/code&gt; table by &lt;code&gt;region&lt;/code&gt;, then show that a query for EU events scans only the EU partition and never touches US storage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;event_id&lt;/th&gt;
&lt;th&gt;region&lt;/th&gt;
&lt;th&gt;user_id&lt;/th&gt;
&lt;th&gt;amount&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;u-eu-1&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;td&gt;u-us-1&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;u-eu-2&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;td&gt;u-us-2&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Region is a partition key, so each region's rows live in their own partition/storage&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;event_id&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;region&lt;/span&gt;   &lt;span class="n"&gt;STRING&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;        &lt;span class="c1"&gt;-- 'EU' | 'US'&lt;/span&gt;
    &lt;span class="n"&gt;user_id&lt;/span&gt;  &lt;span class="n"&gt;STRING&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;amount&lt;/span&gt;   &lt;span class="nb"&gt;DOUBLE&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;PARTITIONED&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- Regional read: the planner prunes to a single partition&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'EU'&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;PARTITIONED BY (region)&lt;/code&gt; lays the four rows into two partitions — &lt;code&gt;region=EU/&lt;/code&gt; holding events 1 and 3, and &lt;code&gt;region=US/&lt;/code&gt; holding events 2 and 4 — each of which can be pinned to region-local storage. When the query filters &lt;code&gt;WHERE region = 'EU'&lt;/code&gt;, the planner performs &lt;em&gt;partition pruning&lt;/em&gt;: it resolves the predicate against the partition key at plan time and lists only the &lt;code&gt;region=EU/&lt;/code&gt; files, so the US partition is never opened. The residency guarantee is therefore structural, not procedural — the EU job has no code path that reads US bytes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;region&lt;/th&gt;
&lt;th&gt;n&lt;/th&gt;
&lt;th&gt;total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;65&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If residency is not visible as a partition key you can filter on, you cannot prove it in a query plan — make &lt;code&gt;region&lt;/code&gt; a partition key first, everything else (pruning, per-region storage, per-region access) follows.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on geo-partitioning for residency
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; An interviewer gives you a global &lt;code&gt;orders&lt;/code&gt; table and a requirement that an EU analyst's session must be physically incapable of scanning non-EU rows. Using partitioning, how do you both store EU rows separately and guarantee an EU-scoped query prunes everything else? Show the DDL and the pruned read.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using partitioning on a region key with partition pruning
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;order_id&lt;/span&gt;  &lt;span class="nb"&gt;BIGINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;region&lt;/span&gt;    &lt;span class="n"&gt;STRING&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;-- residency key: 'EU' | 'US' | 'APAC'&lt;/span&gt;
    &lt;span class="n"&gt;customer&lt;/span&gt;  &lt;span class="n"&gt;STRING&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;amount&lt;/span&gt;    &lt;span class="nb"&gt;DECIMAL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="nb"&gt;TIMESTAMP&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;PARTITIONED&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- Each partition is pinned to region-local storage at load time, e.g.&lt;/span&gt;
&lt;span class="c1"&gt;--   region=EU   -&amp;gt; s3://eu-central-1-orders/ (bucket locked to EU)&lt;/span&gt;
&lt;span class="c1"&gt;--   region=US   -&amp;gt; s3://us-east-1-orders/    (bucket locked to US)&lt;/span&gt;

&lt;span class="c1"&gt;-- EU analyst session runs only region-scoped reads:&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;lifetime_value&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'EU'&lt;/span&gt;            &lt;span class="c1"&gt;-- prunes US and APAC partitions&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;partitions on disk&lt;/th&gt;
&lt;th&gt;predicate&lt;/th&gt;
&lt;th&gt;partitions scanned&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;EU, US, APAC&lt;/td&gt;
&lt;td&gt;(load time)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;EU, US, APAC&lt;/td&gt;
&lt;td&gt;&lt;code&gt;region = 'EU'&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;EU only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;US, APAC&lt;/td&gt;
&lt;td&gt;(never listed)&lt;/td&gt;
&lt;td&gt;pruned&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;aggregate&lt;/td&gt;
&lt;td&gt;grouped in region&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;PARTITIONED BY (region)&lt;/code&gt; writes each region's rows into a separate partition directory that is loaded into a &lt;strong&gt;region-pinned bucket&lt;/strong&gt;, so the physical bytes for &lt;code&gt;US&lt;/code&gt; and &lt;code&gt;APAC&lt;/code&gt; are not even in EU storage.&lt;/li&gt;
&lt;li&gt;The EU query's &lt;code&gt;WHERE region = 'EU'&lt;/code&gt; is a &lt;strong&gt;partition predicate&lt;/strong&gt;; the planner evaluates it against partition metadata and lists only the EU files.&lt;/li&gt;
&lt;li&gt;The US and APAC partitions are &lt;strong&gt;pruned&lt;/strong&gt; — no file handle is opened, so an EU session has no path to non-EU bytes even with a bug in the WHERE clause on non-partition columns.&lt;/li&gt;
&lt;li&gt;The aggregate runs entirely within the EU partition, so both the scan and the result stay in region.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;lifetime_value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;c-eu-ada&lt;/td&gt;
&lt;td&gt;240.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;c-eu-lin&lt;/td&gt;
&lt;td&gt;95.50&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Region as partition key&lt;/strong&gt;&lt;/strong&gt; — promoting &lt;code&gt;region&lt;/code&gt; to a partition key turns residency from a runtime hope into a physical layout, so each region's rows occupy their own files in their own pinned storage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Partition pruning&lt;/strong&gt;&lt;/strong&gt; — because the filter is on the partition key, the planner discards non-EU partitions at plan time; an EU session cannot open US files, which is a stronger guarantee than row-level filtering that still scans everything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Pinned storage per partition&lt;/strong&gt;&lt;/strong&gt; — mapping each partition to a region-locked bucket means the boundary is enforced by the storage layer, not only by the query, so a rogue full-scan still cannot pull a US byte into EU compute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Auditable plan&lt;/strong&gt;&lt;/strong&gt; — the pruned plan is the evidence: you can show an auditor that the EU workload lists only EU partitions, which is exactly the proof residency reviews demand.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — pruning makes the scan O(rows in one region) instead of O(rows globally), so residency and query cost improve together rather than trading off.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Partitioning&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — partitioning&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Partition-key and partition-pruning problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/partitioning" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Sharding&lt;/span&gt;
&lt;span&gt;Topic — sharding&lt;/span&gt;
&lt;strong&gt;Region-shard and data-placement problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/sharding" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. Regional warehouses &amp;amp; replication boundaries
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Give each region its own warehouse and draw a hard line replication may not cross — then let only metadata across
&lt;/h3&gt;

&lt;p&gt;Partitioning keeps rows apart inside one system; the next layer up keeps whole &lt;em&gt;systems&lt;/em&gt; apart. The compliant multi-region pattern is one warehouse deployment per region — a Snowflake account, a BigQuery project/dataset, or a Redshift cluster pinned to that region — with an explicit replication boundary between them that raw regulated data is never allowed to cross. The subtlety interviewers love is that some things &lt;em&gt;should&lt;/em&gt; cross the boundary (schema, lineage, tags, policies) and some things must &lt;em&gt;never&lt;/em&gt; (the actual rows), so the boundary is selective, not a total wall.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One warehouse per region.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Location is immutable and per-deployment.&lt;/strong&gt; A BigQuery dataset's location and a Snowflake account's region are fixed at creation. You do not have one global warehouse with a residency flag; you have N regional warehouses and a routing layer that sends each tenant's queries to the right one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sharding the platform by region.&lt;/strong&gt; Treating each region as an independent shard means a breach, outage, or noisy-neighbour problem in one region is contained to that region's blast radius — the US shard failing does not touch EU data or EU availability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routing by residency key.&lt;/strong&gt; The application resolves a tenant's &lt;code&gt;region&lt;/code&gt; and dispatches to that region's warehouse. The residency key from Section 2 becomes the routing key here — same column, different layer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The replication boundary — what must not cross.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Disable cross-region replication of data.&lt;/strong&gt; Snowflake database replication, BigQuery cross-region dataset copies, Redshift cross-region snapshot copy, and Kafka MirrorMaker are all capable of copying regulated rows across the line. For a residency-bound deployment these are denied by policy, with an allow-list only for non-regulated or already-aggregated data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Geo-fence the streaming layer.&lt;/strong&gt; If Kafka feeds the warehouse, the topics carrying regulated data must live in region-local clusters and MirrorMaker mirroring must be scoped so a EU topic is not mirrored into a US cluster. The stream is as much in scope as the table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backups stay in region.&lt;/strong&gt; Cross-region backup/snapshot copy is the most common accidental breach; the backup destination must be an in-region (or in-jurisdiction) location, verified in policy, not assumed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What may cross — the metadata / data-plane split.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Central control plane, regional data plane.&lt;/strong&gt; Catalog metadata — table names, schemas, column types, lineage, ownership, classification tags, and access policies — is generally &lt;em&gt;not&lt;/em&gt; the regulated payload and can live in a central catalog (e.g. a single governance catalog) that spans regions. The rows themselves stay in the regional data plane.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why the split is safe.&lt;/strong&gt; Knowing that an EU table &lt;code&gt;customers&lt;/code&gt; has a &lt;code&gt;email&lt;/code&gt; column of type &lt;code&gt;string&lt;/code&gt; tagged &lt;code&gt;PII&lt;/code&gt; is metadata; the actual email values are data. Sharing the former globally lets you run one governance model while the latter never leaves the EU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The trap.&lt;/strong&gt; Metadata must be genuinely free of payload — column &lt;em&gt;statistics&lt;/em&gt; like min/max, sample values, or histograms can leak actual data (a min email, a sampled name), so residency-grade catalogs suppress value-bearing statistics for regulated columns.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa47icwwud2e4qkhowkpp.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa47icwwud2e4qkhowkpp.jpeg" alt="Iconographic diagram of regional warehouse deployments — one warehouse account per region behind a hard replication boundary, a central catalog holding only metadata, and cross-region raw replication blocked while control-plane metadata is allowed." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — routing a tenant to its regional warehouse
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The mechanical heart of the pattern is a resolver: given a tenant, return the region, and from the region return the warehouse connection. No query ever names a region literally; it always goes through the resolver, so there is exactly one place that enforces "this tenant's data lives here," and it is trivial to audit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Implement a router that sends a tenant's read to its region's warehouse and refuses if the requested region is outside the tenant's allowed set.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;tenant&lt;/th&gt;
&lt;th&gt;home_region&lt;/th&gt;
&lt;th&gt;requested_region&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;acme-eu&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;globex-us&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;acme-eu&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;WAREHOUSES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;EU&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;snowflake://eu-central-1/acct_eu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;US&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;snowflake://us-east-1/acct_us&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tenant_home_region&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;requested_region&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;requested_region&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;tenant_home_region&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;PermissionError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;residency violation: tenant homed in &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;tenant_home_region&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;may not query &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;requested_region&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;WAREHOUSES&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;requested_region&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;      &lt;span class="c1"&gt;# region-pinned warehouse
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Every request carries the tenant's &lt;code&gt;home_region&lt;/code&gt; (resolved from the residency key) and the region the query wants to touch. The router refuses any request where those differ, so an EU tenant cannot be pointed at the US warehouse even by a mis-built query. When they match, it returns the connection string for that region's &lt;em&gt;own&lt;/em&gt; warehouse account, which is physically pinned to the region — the boundary is enforced at the connection layer before a single row is read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;tenant&lt;/th&gt;
&lt;th&gt;requested_region&lt;/th&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;acme-eu&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;routed to acct_eu&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;globex-us&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;td&gt;routed to acct_us&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;acme-eu&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;td&gt;PermissionError (blocked)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Never let application code name a warehouse region directly — force every query through one resolver keyed on the tenant's residency, so the boundary lives in one auditable function instead of scattered across the codebase.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on replication boundaries
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Product wants a single global dashboard of order counts, but raw EU orders may not leave the EU. An engineer proposes replicating the EU &lt;code&gt;orders&lt;/code&gt; table into the US warehouse to make the join easy. Explain why that breaks residency and design the replication so the boundary holds while the dashboard still works.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using region-local aggregation before any cross-region copy
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- WRONG: replicate raw EU rows to US so a global query can read them&lt;/span&gt;
&lt;span class="c1"&gt;--   CREATE TABLE us.orders_eu_copy AS SELECT * FROM eu.orders;   -- residency breach&lt;/span&gt;

&lt;span class="c1"&gt;-- RIGHT: aggregate inside the EU, replicate only the non-personal summary&lt;/span&gt;
&lt;span class="c1"&gt;-- (runs in the EU warehouse, on EU storage)&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;eu&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders_daily&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;order_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'EU'&lt;/span&gt;        &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;    &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;order_count&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;revenue&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;eu&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;order_date&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;          &lt;span class="c1"&gt;-- no customer, no row-level PII&lt;/span&gt;

&lt;span class="c1"&gt;-- Only orders_daily (counts + revenue, no personal data) is allowed&lt;/span&gt;
&lt;span class="c1"&gt;-- across the replication boundary into the shared reporting layer.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;location&lt;/th&gt;
&lt;th&gt;data shape&lt;/th&gt;
&lt;th&gt;crosses boundary?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;EU warehouse&lt;/td&gt;
&lt;td&gt;raw &lt;code&gt;orders&lt;/code&gt; (PII)&lt;/td&gt;
&lt;td&gt;never&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;EU warehouse&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;orders_daily&lt;/code&gt; aggregate&lt;/td&gt;
&lt;td&gt;eligible&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;boundary policy&lt;/td&gt;
&lt;td&gt;check: no PII columns&lt;/td&gt;
&lt;td&gt;allow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;US / shared layer&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;orders_daily&lt;/code&gt; unioned&lt;/td&gt;
&lt;td&gt;yes (safe)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Replicating raw &lt;code&gt;orders&lt;/code&gt; copies personal rows into US storage governed by US law — that is the exact residency breach, no matter that the query "only wanted counts."&lt;/li&gt;
&lt;li&gt;Instead, the aggregation runs &lt;strong&gt;inside the EU&lt;/strong&gt;, on EU storage, producing &lt;code&gt;orders_daily&lt;/code&gt; that has date, region, count, and revenue — no customer, no personal field.&lt;/li&gt;
&lt;li&gt;The replication boundary policy inspects the object: because &lt;code&gt;orders_daily&lt;/code&gt; carries no PII-tagged columns, it is on the allow-list to cross.&lt;/li&gt;
&lt;li&gt;The shared dashboard unions each region's &lt;code&gt;orders_daily&lt;/code&gt;, so the global view exists while raw rows never left their home region.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_date&lt;/th&gt;
&lt;th&gt;region&lt;/th&gt;
&lt;th&gt;order_count&lt;/th&gt;
&lt;th&gt;revenue&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-03-01&lt;/td&gt;
&lt;td&gt;EU&lt;/td&gt;
&lt;td&gt;1204&lt;/td&gt;
&lt;td&gt;48210.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-03-01&lt;/td&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;td&gt;3391&lt;/td&gt;
&lt;td&gt;121750.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Replication boundary&lt;/strong&gt;&lt;/strong&gt; — the design decision is &lt;em&gt;what&lt;/em&gt; may cross, not merely whether a link exists; raw regulated rows are denied and only derived, non-personal summaries are allowed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Aggregate-in-region&lt;/strong&gt;&lt;/strong&gt; — computing the summary on EU storage means the personal rows are consumed where they live and only the reduced, safe result is a candidate for egress.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Policy on the object, not the intent&lt;/strong&gt;&lt;/strong&gt; — the boundary check inspects the actual columns/tags of what is crossing, so "we only wanted counts" cannot smuggle a &lt;code&gt;SELECT *&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Regional sharding&lt;/strong&gt;&lt;/strong&gt; — each warehouse stays an independent shard, so the shared reporting layer depends on small summaries rather than a fragile cross-region raw replica.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — moving kilobytes of daily aggregates instead of gigabytes of raw orders is both cheaper and the compliant path, so correctness and egress cost align.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Sharding&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — sharding&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Region-shard and boundary-isolation problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/sharding" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Partitioning&lt;/span&gt;
&lt;span&gt;Topic — partitioning&lt;/span&gt;
&lt;strong&gt;Per-region layout and pruning problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/partitioning" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Encryption, key residency &amp;amp; BYOK
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Encrypt every row, but keep the key in the jurisdiction — because whoever holds the key controls the data
&lt;/h3&gt;

&lt;p&gt;Region-pinning and boundaries control &lt;em&gt;where the plaintext sits&lt;/em&gt;. Encryption plus key residency controls &lt;em&gt;who can turn ciphertext back into plaintext&lt;/em&gt;, and that is the lever that actually delivers sovereignty. The design is envelope encryption: each row (or file) is encrypted with a data key, the data key is wrapped by a customer master key, and the master key is held in a key store that resides in the jurisdiction and is controlled by you — so a lawful demand served on a foreign operator yields unreadable bytes, and revoking the key is a one-move "crypto-shred" of the data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Envelope encryption — the mechanism.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Two-tier keys.&lt;/strong&gt; A per-object &lt;strong&gt;data encryption key (DEK)&lt;/strong&gt; encrypts the bytes; a &lt;strong&gt;customer master key (CMK / KEK)&lt;/strong&gt; encrypts the DEK. Only the small wrapped DEK travels with the data; the CMK never leaves the key store.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why two tiers.&lt;/strong&gt; Rotating or revoking one CMK re-controls millions of objects without re-encrypting them all — you re-wrap DEKs, not petabytes. It also means the powerful key can live somewhere much more tightly controlled than the data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where residency bites.&lt;/strong&gt; The CMK's &lt;em&gt;location and control&lt;/em&gt; is the sovereignty control. If the CMK lives in an in-jurisdiction key store you control, the operator holding the ciphertext cannot read it, so extraterritorial reach is blunted by cryptography.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;BYOK, HYOK, and external key stores.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;BYOK (bring your own key).&lt;/strong&gt; You generate the key material and import it into the cloud provider's KMS; you can disable or delete it, which revokes the provider's ability to decrypt. Good, but the key still operates inside the provider's boundary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HYOK / external key store (hold your own key).&lt;/strong&gt; The key material never enters the provider at all — it lives in your own HSM or an external key store (e.g. AWS KMS External Key Store / GCP EKM), and every decrypt makes a call out to your key manager. This is the strongest sovereignty posture short of a fully sovereign cloud, because you can cut access instantly and unilaterally.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Access control on the key is the real control.&lt;/strong&gt; The IAM policy on the CMK decides who can &lt;code&gt;Decrypt&lt;/code&gt;. Sovereignty is enforced by making that policy grant decrypt only to in-jurisdiction principals — the key policy, not the storage location, is where you win or lose.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The multi-region-key foot-gun.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Multi-region KMS keys replicate the key material.&lt;/strong&gt; Cloud "multi-region keys" are a convenience that copies the &lt;em&gt;same&lt;/em&gt; key into several regions so ciphertext is portable. For residency that is precisely wrong: it makes the data decryptable in another jurisdiction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prefer per-region single-region keys.&lt;/strong&gt; Each region gets its own CMK, created and held in that region, so EU ciphertext can only be decrypted by the EU key held under EU control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Crypto-shredding.&lt;/strong&gt; Because access hinges on the key, deleting the region's CMK renders all that region's ciphertext permanently unreadable — a clean, provable deletion that satisfies "right to erasure" without hunting every backup copy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjudiihmbu093xutm7xcg.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjudiihmbu093xutm7xcg.jpeg" alt="Iconographic diagram of envelope encryption with key residency — a plaintext row encrypted by a data key that is wrapped by a region-local customer master key held in an in-region key store, with a revoke-glyph showing crypto-shredding and a warning that multi-region keys replicate." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — per-region CMK wrapping a data key
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The pattern is one CMK per region, each wrapping the DEKs for that region's data. Encrypting an EU row uses the EU CMK; there is no path to decrypt it with the US CMK, because they are different keys held under different control. This makes the key the enforcement point for both residency and sovereignty at once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Encrypt an EU customer row so that only the EU key can decrypt it, and show what a request to decrypt it with the US key produces.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"customer_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"c-eu-1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"region"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"EU"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"email"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ada@example.eu"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;CMK&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;EU&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;eu_key_store&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;US&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;us_key_store&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;   &lt;span class="c1"&gt;# one master key per region, held in-region
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;encrypt_row&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;region&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;dek&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;generate_data_key&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                     &lt;span class="c1"&gt;# random per-object DEK
&lt;/span&gt;    &lt;span class="n"&gt;ciphertext&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;aes_encrypt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dek&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;            &lt;span class="c1"&gt;# encrypt the bytes
&lt;/span&gt;    &lt;span class="n"&gt;wrapped_dek&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;CMK&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;wrap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dek&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;           &lt;span class="c1"&gt;# CMK never leaves the store
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wrapped_dek&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;wrapped_dek&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ciphertext&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decrypt_row&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;use_region&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;dek&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;CMK&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;use_region&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;unwrap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wrapped_dek&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;   &lt;span class="c1"&gt;# fails unless region matches
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;aes_decrypt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dek&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;encrypt_row&lt;/code&gt; picks the CMK for the row's &lt;code&gt;region&lt;/code&gt;, generates a fresh DEK, encrypts the bytes with the DEK, and asks the EU CMK to wrap the DEK — the CMK itself stays inside the EU key store. To read the row you must unwrap the DEK with the &lt;em&gt;same&lt;/em&gt; CMK that wrapped it. Calling &lt;code&gt;decrypt_row(obj, "US")&lt;/code&gt; asks the US key store to unwrap a DEK wrapped by the EU key; the US CMK is a different key and the unwrap fails, so the ciphertext is inert outside the EU.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;operation&lt;/th&gt;
&lt;th&gt;key used&lt;/th&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;encrypt EU row&lt;/td&gt;
&lt;td&gt;EU CMK&lt;/td&gt;
&lt;td&gt;wrapped DEK + ciphertext&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;decrypt with EU CMK&lt;/td&gt;
&lt;td&gt;EU CMK&lt;/td&gt;
&lt;td&gt;plaintext row&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;decrypt with US CMK&lt;/td&gt;
&lt;td&gt;US CMK&lt;/td&gt;
&lt;td&gt;unwrap fails — ciphertext stays opaque&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; One CMK per region, single-region, held in-jurisdiction — never a multi-region key — so the only way to read a region's data is a key that region controls.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on key residency and access revocation
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Regulators order you to make an EU dataset immediately unreadable to a specific processor, across the live table &lt;em&gt;and every backup&lt;/em&gt;, without physically deleting terabytes. How does an in-region key design let you do this in one operation, and what access-control setting makes it enforceable?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using access-control on a region-local CMK (crypto-shred)
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- All EU data is encrypted with a DEK wrapped by the EU CMK.&lt;/span&gt;
&lt;span class="c1"&gt;-- The table (and every backup) stores only ciphertext + wrapped DEK.&lt;/span&gt;

&lt;span class="c1"&gt;-- Access to decrypt is a KEY POLICY, not a table grant:&lt;/span&gt;
&lt;span class="c1"&gt;--   grant kms:Decrypt on cmk_eu to  &amp;lt;in-region service principals&amp;gt;&lt;/span&gt;
&lt;span class="c1"&gt;--   deny  kms:Decrypt on cmk_eu to  &amp;lt;the ordered-out processor&amp;gt;&lt;/span&gt;

&lt;span class="c1"&gt;-- To make the dataset unreadable to that processor RIGHT NOW:&lt;/span&gt;
&lt;span class="k"&gt;REVOKE&lt;/span&gt; &lt;span class="n"&gt;DECRYPT&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt; &lt;span class="n"&gt;cmk_eu&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="k"&gt;ROLE&lt;/span&gt; &lt;span class="n"&gt;processor_x&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;-- single control-plane op&lt;/span&gt;

&lt;span class="c1"&gt;-- To crypto-shred the WHOLE dataset (live + all backups) irreversibly:&lt;/span&gt;
&lt;span class="c1"&gt;--   schedule deletion of cmk_eu  -&amp;gt;  every ciphertext ever wrapped by it&lt;/span&gt;
&lt;span class="c1"&gt;--   becomes permanently undecryptable, no matter which backup holds it.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;action&lt;/th&gt;
&lt;th&gt;scope affected&lt;/th&gt;
&lt;th&gt;data physically deleted?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;revoke &lt;code&gt;Decrypt&lt;/code&gt; from processor_x&lt;/td&gt;
&lt;td&gt;live + backups&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;processor_x reads table&lt;/td&gt;
&lt;td&gt;ciphertext only&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;(erasure) schedule CMK deletion&lt;/td&gt;
&lt;td&gt;every wrapped DEK&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;any future read, any copy&lt;/td&gt;
&lt;td&gt;unwrap fails&lt;/td&gt;
&lt;td&gt;effectively yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Because decryption requires &lt;code&gt;kms:Decrypt&lt;/code&gt; on the EU CMK, revoking that permission from &lt;code&gt;processor_x&lt;/code&gt; instantly removes their ability to read — for the live table &lt;em&gt;and&lt;/em&gt; every backup, since all copies share the same wrapping key.&lt;/li&gt;
&lt;li&gt;The processor can still fetch bytes, but without the key those bytes are ciphertext, so the revocation is a single control-plane change with immediate, global effect.&lt;/li&gt;
&lt;li&gt;For full erasure, scheduling deletion of the CMK is one operation whose reach is every object the key ever wrapped, wherever it is stored.&lt;/li&gt;
&lt;li&gt;After the key is gone, no principal — not even you — can unwrap those DEKs, so the data is provably unreadable without hunting down each physical copy.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;target&lt;/th&gt;
&lt;th&gt;before revoke&lt;/th&gt;
&lt;th&gt;after revoke&lt;/th&gt;
&lt;th&gt;after CMK deletion&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;processor_x reads&lt;/td&gt;
&lt;td&gt;plaintext&lt;/td&gt;
&lt;td&gt;denied (ciphertext)&lt;/td&gt;
&lt;td&gt;denied (ciphertext)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;everyone reads&lt;/td&gt;
&lt;td&gt;plaintext&lt;/td&gt;
&lt;td&gt;plaintext&lt;/td&gt;
&lt;td&gt;unreadable (shredded)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Key policy as access control&lt;/strong&gt;&lt;/strong&gt; — decryption is gated by &lt;code&gt;kms:Decrypt&lt;/code&gt; on the CMK, so who-can-read is a key-policy decision that applies uniformly to the live table and every derived copy at once.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Envelope indirection&lt;/strong&gt;&lt;/strong&gt; — because every object's DEK is wrapped by the one region CMK, a single key action re-governs or destroys millions of rows without touching the rows themselves.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Crypto-shredding&lt;/strong&gt;&lt;/strong&gt; — deleting the region key makes all ciphertext it wrapped permanently opaque, giving a provable, backup-inclusive "right to erasure" that physical deletion sweeps can never guarantee.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;In-region, single-region key&lt;/strong&gt;&lt;/strong&gt; — holding one CMK per region under in-jurisdiction control is what makes both the revoke and the shred lawful, unilateral, and outside a foreign operator's reach.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the operation is O(1) key actions rather than O(bytes) re-encryption or deletion, so a compliance mandate becomes a single reversible-then-irreversible control-plane step.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Access control&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — access-control&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Key-policy and access-revocation problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Partitioning&lt;/span&gt;
&lt;span&gt;Topic — partitioning&lt;/span&gt;
&lt;strong&gt;Per-region key and data-placement problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/partitioning" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Cross-region aggregate-only egress
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Keep the rows home and let only insight travel — spokes aggregate, a k-anonymity gate filters, the hub unions sums
&lt;/h3&gt;

&lt;p&gt;The final pattern answers the question every business eventually asks: "we need a global view, but the data cannot leave its region — now what?" The answer is aggregate-only egress in a hub-and-spoke shape. Each regional spoke computes aggregates locally, a privacy gate (a minimum group size, a.k.a. k-anonymity) drops any aggregate small enough to re-identify an individual, and only the surviving, safe summaries cross to a central hub that unions them. Raw rows never move; the metadata/control plane is central; the data plane stays regional.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Aggregate-only, computed in region.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reduce before you move.&lt;/strong&gt; The spoke runs &lt;code&gt;GROUP BY&lt;/code&gt; on region-local storage, so what leaves is &lt;code&gt;COUNT&lt;/code&gt;/&lt;code&gt;SUM&lt;/code&gt;/&lt;code&gt;AVG&lt;/code&gt; per group — not the underlying rows. The reduction is the compliance boundary: individuals are gone by the time anything crosses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Suppress small groups (k-anonymity).&lt;/strong&gt; An aggregate over a tiny group can re-identify a person (a &lt;code&gt;SUM(salary)&lt;/code&gt; for one employee &lt;em&gt;is&lt;/em&gt; that salary). A &lt;code&gt;HAVING count(*) &amp;gt;= k&lt;/code&gt; threshold suppresses groups below k so no aggregate leaks an individual. Common k values are 5–20 depending on sensitivity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Beware differencing attacks.&lt;/strong&gt; Two overlapping aggregates can be subtracted to reveal a suppressed cell; robust designs add consistent suppression rules (or noise / differential privacy) so the hub cannot reconstruct a small group by arithmetic.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Hub-and-spoke topology.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Spokes = regional data planes.&lt;/strong&gt; Each region is a full, independent shard holding its raw data and doing its own aggregation — the same regional-warehouse shards from Section 3.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hub = a thin global layer.&lt;/strong&gt; The hub stores only the safe aggregates and unions them into global metrics. It holds no raw regional data, so the hub's own jurisdiction is not a residency problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The control plane is separate from the data plane.&lt;/strong&gt; Schema, job definitions, lineage, and policies (metadata) are managed centrally and pushed to spokes; the data (rows) never flow the other way. This is the same metadata/data split as Section 3, applied to analytics.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;When even aggregates are not enough — sovereign cloud.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sovereign-cloud offerings.&lt;/strong&gt; When regulation demands that even the &lt;em&gt;operator&lt;/em&gt; be in-jurisdiction (in-country support staff, in-country legal entity, isolated from foreign parent control), providers offer sovereign-cloud options — for example AWS European Sovereign Cloud, Microsoft Cloud for Sovereignty, and Google Sovereign Controls — plus national efforts like Gaia-X.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confidential computing.&lt;/strong&gt; Encrypting data &lt;em&gt;in use&lt;/em&gt; (enclaves / trusted execution) closes the last gap where plaintext exists in memory during processing, so even the operator's admins cannot read it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Choose the least-heavy control that clears the bar.&lt;/strong&gt; Region-pinning handles residency; customer-held keys handle most sovereignty; sovereign-cloud and confidential computing are for the strictest mandates. Reaching for the heaviest control everywhere is expensive and slow — match the control to the legal requirement.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4imkejlgosg5q0qz13hb.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4imkejlgosg5q0qz13hb.jpeg" alt="Iconographic hub-and-spoke diagram — regional spokes compute local aggregates behind a k-anonymity gate, only small-count-suppressed sums cross to a central hub that unions them, while raw rows and small groups stay in region." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — a spoke aggregate that suppresses small groups
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday building block is a spoke query that aggregates and applies the k-anonymity &lt;code&gt;HAVING&lt;/code&gt; filter in the same statement, so the object that leaves the region is &lt;em&gt;already&lt;/em&gt; both reduced and suppressed. Nothing downstream has to remember to re-apply the gate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; In the EU spoke, produce per-city order counts that may leave the region, suppressing any city with fewer than 5 customers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;city&lt;/th&gt;
&lt;th&gt;customers&lt;/th&gt;
&lt;th&gt;orders&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Berlin&lt;/td&gt;
&lt;td&gt;900&lt;/td&gt;
&lt;td&gt;3400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Munich&lt;/td&gt;
&lt;td&gt;120&lt;/td&gt;
&lt;td&gt;510&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aland (tiny)&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Runs in the EU spoke, on EU storage. Output is egress-eligible.&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;city&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                    &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;eu&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;city&lt;/span&gt;
&lt;span class="k"&gt;HAVING&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;-- k-anonymity: drop tiny groups&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The &lt;code&gt;GROUP BY city&lt;/code&gt; reduces raw orders to one row per city — individuals are already gone. The &lt;code&gt;HAVING COUNT(DISTINCT customer_id) &amp;gt;= 5&lt;/code&gt; clause drops any city whose group is smaller than k=5, so "Aland" with 3 customers is suppressed and never appears in the output. What remains is a small, non-identifying summary that is safe to send to the hub; the raw &lt;code&gt;eu.orders&lt;/code&gt; rows stay in the EU.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;city&lt;/th&gt;
&lt;th&gt;customers&lt;/th&gt;
&lt;th&gt;orders&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Berlin&lt;/td&gt;
&lt;td&gt;900&lt;/td&gt;
&lt;td&gt;3400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Munich&lt;/td&gt;
&lt;td&gt;120&lt;/td&gt;
&lt;td&gt;510&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Apply the k-anonymity &lt;code&gt;HAVING&lt;/code&gt; filter in the same query that aggregates, in-region — so the artifact that crosses the boundary is provably suppressed and no downstream step can forget the gate.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on aggregate-only cross-region reporting
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You must build a global "revenue by product category" report from EU and US spokes without any raw row leaving its region, and a category with fewer than 10 customers in a region must be suppressed. Design the spoke query and the hub union.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using in-region aggregation, a k-anonymity gate, and a hub union
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- SPOKE (runs identically in EU and US, each on its own region-local storage)&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;}.&lt;/span&gt;&lt;span class="n"&gt;category_daily&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="s1"&gt;'EU'&lt;/span&gt;                        &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;     &lt;span class="c1"&gt;-- literal per spoke&lt;/span&gt;
    &lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                 &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;revenue&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;eu&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;category&lt;/span&gt;
&lt;span class="k"&gt;HAVING&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;      &lt;span class="c1"&gt;-- suppress &amp;lt; 10 customers&lt;/span&gt;

&lt;span class="c1"&gt;-- HUB (thin global layer; stores only the safe aggregates, no raw rows)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;revenue&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;global_revenue&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;eu&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category_daily&lt;/span&gt;            &lt;span class="c1"&gt;-- crossed the boundary: safe&lt;/span&gt;
    &lt;span class="k"&gt;UNION&lt;/span&gt; &lt;span class="k"&gt;ALL&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;us&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category_daily&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;location&lt;/th&gt;
&lt;th&gt;rows&lt;/th&gt;
&lt;th&gt;privacy gate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;EU spoke&lt;/td&gt;
&lt;td&gt;raw &lt;code&gt;orders&lt;/code&gt; (PII)&lt;/td&gt;
&lt;td&gt;stays in EU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;EU spoke&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;category_daily&lt;/code&gt; (&amp;gt;=10)&lt;/td&gt;
&lt;td&gt;small cats suppressed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;US spoke&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;category_daily&lt;/code&gt; (&amp;gt;=10)&lt;/td&gt;
&lt;td&gt;small cats suppressed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;hub&lt;/td&gt;
&lt;td&gt;union + re-aggregate&lt;/td&gt;
&lt;td&gt;only safe sums&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Each spoke aggregates its own raw orders &lt;strong&gt;in region&lt;/strong&gt;, so the personal rows never move; only &lt;code&gt;category&lt;/code&gt;, &lt;code&gt;customers&lt;/code&gt;, and &lt;code&gt;revenue&lt;/code&gt; are produced.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;HAVING ... &amp;gt;= 10&lt;/code&gt; gate suppresses any category with fewer than 10 customers in that region, so no small-group aggregate can re-identify a person before egress.&lt;/li&gt;
&lt;li&gt;The two suppressed summaries are the &lt;em&gt;only&lt;/em&gt; objects that cross the replication boundary into the hub.&lt;/li&gt;
&lt;li&gt;The hub &lt;code&gt;UNION ALL&lt;/code&gt;s the regional summaries and re-aggregates to a global figure; it holds no raw data, so its own location is not a residency concern.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;category&lt;/th&gt;
&lt;th&gt;global_revenue&lt;/th&gt;
&lt;th&gt;customers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;electronics&lt;/td&gt;
&lt;td&gt;512000.00&lt;/td&gt;
&lt;td&gt;4100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;grocery&lt;/td&gt;
&lt;td&gt;288500.00&lt;/td&gt;
&lt;td&gt;9600&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Aggregate-in-region&lt;/strong&gt;&lt;/strong&gt; — reducing raw rows to group sums on region-local storage means individuals are eliminated before anything is eligible to move, so egress carries insight, not people.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;k-anonymity gate&lt;/strong&gt;&lt;/strong&gt; — the &lt;code&gt;HAVING count &amp;gt;= k&lt;/code&gt; filter guarantees every crossing aggregate covers enough individuals to be non-identifying, converting "aggregate" from "usually safe" into "provably suppressed."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Hub-and-spoke&lt;/strong&gt;&lt;/strong&gt; — spokes are independent regional shards holding the data; the hub is a thin union layer holding only safe summaries, so no single place concentrates cross-region raw rows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Metadata/data split&lt;/strong&gt;&lt;/strong&gt; — the schema and job run centrally while the rows stay regional, giving one global report definition without one global data pile.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the hub moves and stores O(groups) aggregates instead of O(rows) raw data, so the compliant design is also the cheaper and faster one to run globally.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Partitioning&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — partitioning&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;In-region aggregation and layout problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/partitioning" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Access control&lt;/span&gt;
&lt;span&gt;Topic — access-control&lt;/span&gt;
&lt;strong&gt;Egress-policy and suppression problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — residency recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Pin a bucket region and block replication.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;S3 bucket lives in one region for life; deny replication out of the boundary
- create bucket in eu-central-1
- bucket policy / SCP: DENY s3:PutBucketReplication to non-EU regions
- DENY cross-region snapshot/backup copy for this bucket
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Geo-partition on a region key.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event_id&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt; &lt;span class="n"&gt;STRING&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="n"&gt;STRING&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;PARTITIONED&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;          &lt;span class="c1"&gt;-- region=EU/ , region=US/  in pinned storage&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'EU'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;-- prunes US partitions&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Route a tenant to its regional warehouse.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;home_region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;requested_region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;warehouses&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;requested_region&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;home_region&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;PermissionError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;residency violation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;warehouses&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;requested_region&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;     &lt;span class="c1"&gt;# region-pinned account
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Per-region CMK envelope encryption.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;dek&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;generate_data_key&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;ct&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;aes_encrypt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dek&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;wrapped&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;CMK&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]].&lt;/span&gt;&lt;span class="nf"&gt;wrap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dek&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="c1"&gt;# one single-region CMK per region
&lt;/span&gt;&lt;span class="n"&gt;dek_out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;CMK&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]].&lt;/span&gt;&lt;span class="nf"&gt;unwrap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;wrapped&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# decrypt needs the SAME region CMK; delete it -&amp;gt; crypto-shred
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;k-anonymity aggregate for egress.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;city&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;revenue&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;eu&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;city&lt;/span&gt;
&lt;span class="k"&gt;HAVING&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;-- suppress small groups, then egress&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Residency decision table.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requirement&lt;/th&gt;
&lt;th&gt;Control&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bytes must stay in region&lt;/td&gt;
&lt;td&gt;Region-pinned storage + &lt;code&gt;PARTITION BY region&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One region must not read another&lt;/td&gt;
&lt;td&gt;Partition pruning + per-region warehouse routing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No raw rows across the boundary&lt;/td&gt;
&lt;td&gt;Aggregate-in-region + object-level egress policy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operator must not be able to read&lt;/td&gt;
&lt;td&gt;BYOK / HYOK, in-region single-region CMK&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Even the operator must be in-jurisdiction&lt;/td&gt;
&lt;td&gt;Sovereign cloud + confidential computing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Global view without moving rows&lt;/td&gt;
&lt;td&gt;Hub-and-spoke, aggregate-only + k-anonymity&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is the difference between data residency and data sovereignty?
&lt;/h3&gt;

&lt;p&gt;Data residency is a statement about &lt;strong&gt;geography&lt;/strong&gt; — that the physical bytes of a dataset are stored and processed within a named region and its copies (replicas, backups, temp files, caches) do not leave it. Data sovereignty is a statement about &lt;strong&gt;jurisdiction&lt;/strong&gt; — whose laws can compel access to the data regardless of where it sits. You can satisfy residency (store EU data in the EU) while failing sovereignty (a foreign-headquartered operator remains legally compellable), which is why sovereignty controls like in-region operators and customer-held keys exist on top of region-pinning.&lt;/p&gt;

&lt;h3&gt;
  
  
  How is data residency different from disaster recovery?
&lt;/h3&gt;

&lt;p&gt;Disaster recovery deliberately creates &lt;strong&gt;copies&lt;/strong&gt; across failure domains so a regional outage does not cause data loss or downtime; residency deliberately &lt;strong&gt;bounds&lt;/strong&gt; where copies may exist so data stays in one jurisdiction. The two pull in opposite directions, and cross-region DR replication is the single most common way a residency guarantee is silently broken. The fix is to keep DR inside the residency boundary — a second in-jurisdiction zone or region — rather than backing up to whichever region is cheapest.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you keep data from crossing a region boundary?
&lt;/h3&gt;

&lt;p&gt;Make &lt;code&gt;region&lt;/code&gt; a partition key so each region's rows live in their own partitions pinned to region-local storage, deploy one warehouse per region, and explicitly deny cross-region replication, snapshot copy, and stream mirroring for regulated data. Route every query through a resolver keyed on the tenant's residency so no code can point an EU session at US storage. When a global view is needed, aggregate inside each region and let only non-personal summaries cross the boundary, never raw rows.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is BYOK and why does key residency matter?
&lt;/h3&gt;

&lt;p&gt;BYOK (bring your own key) means you generate the encryption key material and import it into the provider's key service, keeping the power to disable or delete it; HYOK/external key stores go further and keep the key entirely in your own HSM so the provider never holds it. Key residency matters because whoever controls the key controls the data — envelope encryption wraps each object's data key with your region-local customer master key, so ciphertext stored under a foreign operator is unreadable without a key that stays in your jurisdiction, and revoking or deleting that key crypto-shreds the data across the live table and every backup at once.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can metadata leave the region if the data cannot?
&lt;/h3&gt;

&lt;p&gt;Usually yes — schema, column types, lineage, ownership, and classification tags are metadata, not the regulated payload, so a central catalog can span regions while the rows stay in the regional data plane. The important caveat is that value-bearing statistics (min/max, sampled values, histograms) can leak actual data, so a residency-grade catalog suppresses those for regulated columns. This metadata/data-plane split is what lets you run one global governance model without concentrating regulated data in one place.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is a sovereign cloud?
&lt;/h3&gt;

&lt;p&gt;A sovereign cloud is an offering designed so that not just the data but the &lt;strong&gt;operation&lt;/strong&gt; stays in-jurisdiction — in-country legal entity, in-country support staff, and isolation from foreign-parent control — for mandates where region-pinning and customer keys are not enough. Examples include AWS European Sovereign Cloud, Microsoft Cloud for Sovereignty, and Google Sovereign Controls, alongside national initiatives like Gaia-X. Combined with confidential computing (encrypting data in use inside enclaves), it closes the gap where even the operator's administrators might otherwise access plaintext.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every residency idea above, from geo-partitioning on a region key to per-region warehouse routing, in-region CMK crypto-shredding, and k-anonymity aggregate egress, maps to a hands-on practice room where you build the layout against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you keep EU rows from ever reaching US compute?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/partitioning" rel="noopener noreferrer"&gt;Practice partitioning problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/sharding" rel="noopener noreferrer"&gt;Region-shard drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>GDPR vs CCPA vs India DPDP: Building Multi-Jurisdiction Privacy Pipelines</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Fri, 25 Sep 2026 13:44:49 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/gdpr-vs-ccpa-vs-india-dpdp-building-multi-jurisdiction-privacy-pipelines-4g3f</link>
      <guid>https://dev.to/gowthampotureddi/gdpr-vs-ccpa-vs-india-dpdp-building-multi-jurisdiction-privacy-pipelines-4g3f</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;multi-jurisdiction privacy pipelines&lt;/code&gt;&lt;/strong&gt; are what you build when the same row of personal data must be legal in three places at once — an EU customer, a California consumer, and an Indian data principal all land in the same warehouse, and each is governed by a different statute with a different vocabulary, a different legal trigger for processing, and a different deadline for honouring a deletion request. The naive answer is to fork the platform per region. The engineering answer is to parameterise one platform by jurisdiction, so that consent, policy, subject rights, and de-identification are data-driven rather than region-specific code.&lt;/p&gt;

&lt;p&gt;This guide is a contrast, not a single-law walkthrough. The EU's GDPR, California's CCPA (as amended by the CPRA), and India's Digital Personal Data Protection Act of 2023 agree on the goal — give individuals control over their data — and disagree on almost every mechanism: opt-in consent versus opt-out of sale, a "data subject" versus a "consumer" versus a "data principal", a one-month response window versus forty-five days. A data engineer who hard-codes GDPR everywhere ships a system that is simultaneously over-compliant in California and non-compliant in India. The four ideas an interviewer will actually probe — a unified consent and subject registry, per-jurisdiction policy tags with purpose limitation, the access and erasure pipelines, and tokenization versus pseudonymization — are the load-bearing walls of a platform that satisfies all three at once.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frf4i0ysc47kuq6rgzf8e.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frf4i0ysc47kuq6rgzf8e.jpeg" alt="PipeCode blog header for multi-jurisdiction privacy pipelines — bold white headline 'GDPR · CCPA · DPDP' with subtitle 'registry · policy tags · erasure · tokenization' and a stylised three-region-to-one-platform scene on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;access-control practice library →&lt;/a&gt;, rehearse de-identification on the &lt;a href="https://pipecode.ai/explore/practice/topic/tokenization" rel="noopener noreferrer"&gt;tokenization practice set →&lt;/a&gt;, and harden your validation logic on the &lt;a href="https://pipecode.ai/explore/practice/topic/data-quality" rel="noopener noreferrer"&gt;data-quality practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Privacy as a data-engineering problem, not a legal memo&lt;/li&gt;
&lt;li&gt;The unified consent &amp;amp; subject registry&lt;/li&gt;
&lt;li&gt;Policy tags, data mapping &amp;amp; purpose limitation&lt;/li&gt;
&lt;li&gt;Subject-rights pipelines — access &amp;amp; erasure tombstones&lt;/li&gt;
&lt;li&gt;Tokenization &amp;amp; pseudonymization&lt;/li&gt;
&lt;li&gt;Cheat sheet — multi-jurisdiction privacy recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Privacy as a data-engineering problem, not a legal memo
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Three statutes, one physical platform — the differences are the specification, not trivia
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;you cannot fork the warehouse per country, so every difference between GDPR, CCPA, and DPDP must become a column, a tag, or a policy the pipeline reads at runtime&lt;/strong&gt;. A privacy pipeline is not a lawyer's summary rendered in a wiki; it is the set of data structures and jobs that make "is this processing allowed?" and "erase this person" answerable by code. To build it you have to know precisely where the three regimes diverge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The vocabulary maps, but does not match.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The individual.&lt;/strong&gt; GDPR calls them a &lt;strong&gt;data subject&lt;/strong&gt;; CCPA calls them a &lt;strong&gt;consumer&lt;/strong&gt; (a California resident); DPDP calls them a &lt;strong&gt;data principal&lt;/strong&gt;. Same human, three keys — which is exactly why you need one canonical &lt;code&gt;subject_id&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The accountable party.&lt;/strong&gt; GDPR has a &lt;strong&gt;controller&lt;/strong&gt; (decides purpose and means) and a &lt;strong&gt;processor&lt;/strong&gt;; CCPA has a &lt;strong&gt;business&lt;/strong&gt; and a &lt;strong&gt;service provider&lt;/strong&gt;; DPDP has a &lt;strong&gt;data fiduciary&lt;/strong&gt; and a &lt;strong&gt;data processor&lt;/strong&gt;, plus the higher-duty &lt;strong&gt;Significant Data Fiduciary (SDF)&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The oversight role.&lt;/strong&gt; GDPR can require a &lt;strong&gt;Data Protection Officer (DPO)&lt;/strong&gt;; DPDP requires an SDF to appoint a &lt;strong&gt;DPO based in India&lt;/strong&gt;; CCPA has no DPO mandate but the CPRA created the California Privacy Protection Agency (CPPA) as a regulator.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The legal trigger for processing is where they diverge most.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GDPR — a lawful basis, opt-in by default.&lt;/strong&gt; You may only process personal data if you can name one of six lawful bases (Article 6): consent, contract, legal obligation, vital interests, public task, or legitimate interests. Consent, when used, must be freely given, specific, informed, and unambiguous.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CCPA/CPRA — notice and opt-out, not consent.&lt;/strong&gt; California does not require a lawful basis to collect. It requires &lt;strong&gt;notice at collection&lt;/strong&gt; and gives the consumer a &lt;strong&gt;right to opt out of the sale or sharing&lt;/strong&gt; of personal information (the "Do Not Sell or Share" link and the Global Privacy Control signal). Opt-in consent is only mandated for minors and for "sensitive personal information" limits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DPDP — consent or a "legitimate use", plus itemised notice.&lt;/strong&gt; India's default is &lt;strong&gt;consent&lt;/strong&gt; (free, specific, informed, unconditional, unambiguous, with a clear affirmative action), backed by a plain-language &lt;strong&gt;notice&lt;/strong&gt; that must be available in English and the scheduled Indian languages. A short list of "legitimate uses" (e.g. the principal voluntarily provides data for a requested service) can substitute.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The subject rights rhyme, but the deadlines and shapes differ.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Access.&lt;/strong&gt; GDPR's Data Subject Access Request (DSAR, Article 15) → &lt;strong&gt;1 month&lt;/strong&gt;. CCPA's right to know → &lt;strong&gt;45 days&lt;/strong&gt; (extendable to 90). DPDP's right to access a summary of processing → within a period the rules set, via the fiduciary's grievance mechanism.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Erasure.&lt;/strong&gt; GDPR's right to erasure ("right to be forgotten", Article 17) is broad. CCPA's right to delete has enumerated exceptions. DPDP ties erasure to &lt;strong&gt;withdrawal of consent&lt;/strong&gt; and to the purpose being fulfilled — there is no standalone "right to be forgotten" clause, but withdrawal must be as easy as giving consent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Correction and portability.&lt;/strong&gt; GDPR grants rectification and portability; CPRA added a right to correct; DPDP grants correction and completion. Portability is a first-class GDPR right and absent from DPDP.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you say &lt;strong&gt;"I parameterise the platform by jurisdiction"&lt;/strong&gt; rather than "I build a GDPR system"? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you distinguish &lt;strong&gt;opt-in consent (GDPR/DPDP) from opt-out of sale (CCPA)&lt;/strong&gt; without prompting? — the single most common trap.&lt;/li&gt;
&lt;li&gt;Do you treat &lt;strong&gt;erasure as propagation across every downstream copy&lt;/strong&gt;, not a &lt;code&gt;DELETE&lt;/code&gt; on one table? — the whole engineering point.&lt;/li&gt;
&lt;li&gt;Do you know that &lt;strong&gt;DPDP has no portability right and CCPA needs no lawful basis&lt;/strong&gt;, so a one-size GDPR build is wrong in two directions? — depth signal.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — the same person under three regimes
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; Before any code, internalise that one physical record carries three legal states simultaneously. A single user, Priya, who is an EU resident who later moved to California and holds Indian citizenship, could plausibly trigger all three regimes depending on where the data was collected and where she resides. The platform must answer "what may I do with row X?" differently per jurisdiction, from the &lt;em&gt;same&lt;/em&gt; stored row.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; For one stored profile row, list what each regime requires before you may use it for marketing analytics, and what a deletion request means under each.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;attribute&lt;/th&gt;
&lt;th&gt;value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;subject_id&lt;/td&gt;
&lt;td&gt;&lt;code&gt;s_88f1&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;email&lt;/td&gt;
&lt;td&gt;&lt;code&gt;priya@example.com&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;collected_in&lt;/td&gt;
&lt;td&gt;EU (2023), then CA (2025)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;purpose_requested&lt;/td&gt;
&lt;td&gt;marketing analytics&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;REGIMES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;   &lt;span class="c1"&gt;# a jurisdiction is a policy object, not a country string in a WHERE clause
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GDPR&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;needs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lawful_basis&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;marketing_ok_if&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;consent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;erasure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;broad&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;access_sla_days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CCPA&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;needs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notice_at_collection&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;marketing_ok_if&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;not_opted_out&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;erasure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;with_exceptions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;access_sla_days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;45&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DPDP&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;needs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;consent_or_legitimate_use&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;marketing_ok_if&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;consent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;erasure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;on_withdrawal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;access_sla_days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Under &lt;strong&gt;GDPR&lt;/strong&gt;, marketing analytics on personal data needs a lawful basis; for marketing that is almost always &lt;em&gt;consent&lt;/em&gt;, so you may proceed only if Priya opted in, and a deletion request is a broad erasure. Under &lt;strong&gt;CCPA&lt;/strong&gt;, you needed no consent to collect — only notice — so marketing is allowed &lt;em&gt;unless&lt;/em&gt; she opted out of sale/sharing, and deletion is honoured subject to statutory exceptions. Under &lt;strong&gt;DPDP&lt;/strong&gt;, marketing requires consent (or a listed legitimate use) plus a valid notice id, and "deletion" is triggered when she withdraws that consent. Same row, three gates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;regime&lt;/th&gt;
&lt;th&gt;may run marketing analytics if…&lt;/th&gt;
&lt;th&gt;deletion means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GDPR&lt;/td&gt;
&lt;td&gt;consent on record (opt-in)&lt;/td&gt;
&lt;td&gt;broad erasure across copies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CCPA&lt;/td&gt;
&lt;td&gt;consumer has not opted out&lt;/td&gt;
&lt;td&gt;delete with enumerated exceptions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DPDP&lt;/td&gt;
&lt;td&gt;consent + valid notice id&lt;/td&gt;
&lt;td&gt;erase on consent withdrawal&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If a single answer ("delete the user") is the same across all three regimes in your design, you have almost certainly hard-coded one law and mis-modelled the other two.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. The unified consent &amp;amp; subject registry
&lt;/h2&gt;

&lt;h3&gt;
  
  
  One canonical subject_id and an append-only consent ledger are the foundation everything else stands on
&lt;/h3&gt;

&lt;p&gt;The feature that makes a multi-jurisdiction platform tractable is that &lt;strong&gt;every regional identity resolves to one canonical &lt;code&gt;subject_id&lt;/code&gt;, and every legality fact about that subject is an event in an append-only ledger&lt;/strong&gt;. Without the single id you cannot honour an erasure request that spans an EU login and a California account; without the append-only ledger you cannot prove &lt;em&gt;what consent existed at the moment a job ran&lt;/em&gt;, which is the question a regulator actually asks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Identity resolution first.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One subject, many keys.&lt;/strong&gt; An EU customer id, a California consumer id, and an Indian principal id may all be the same human. A resolution step (deterministic on verified email/phone, probabilistic only as a fallback) maps them to one &lt;code&gt;subject_id&lt;/code&gt;. This id, never raw PII, is the join key across the platform.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why not just email.&lt;/strong&gt; Emails change, are shared, and are themselves PII. The &lt;code&gt;subject_id&lt;/code&gt; is an opaque surrogate so that analytics can join on identity without touching a regulated attribute.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The consent ledger is event-sourced, not a mutable flag.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Per-jurisdiction, per-purpose rows.&lt;/strong&gt; A ledger row records &lt;code&gt;(subject_id, jurisdiction, purpose, basis, state, notice_id, ts, source)&lt;/code&gt;. GDPR rows carry a &lt;code&gt;lawful_basis&lt;/code&gt;; CCPA rows carry an &lt;code&gt;opt_out&lt;/code&gt; state for sale/sharing; DPDP rows carry &lt;code&gt;consent&lt;/code&gt; plus the &lt;code&gt;notice_id&lt;/code&gt; the principal saw.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Withdrawal is a new event.&lt;/strong&gt; You never &lt;code&gt;UPDATE consent=false&lt;/code&gt;. You append a &lt;code&gt;withdrawn&lt;/code&gt; event with a timestamp. The current state is a fold over history, so the ledger can answer "was consent valid at 14:03 last Tuesday?" — required for audits and for defending a past processing decision.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Versioned notices.&lt;/strong&gt; Because DPDP (and GDPR transparency) tie validity to the exact notice text shown, the &lt;code&gt;notice_id&lt;/code&gt; references an immutable, versioned notice record. Re-consent is required when a material purpose changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The gate that reads the ledger.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;is_allowed(subject_id, purpose, jurisdiction)&lt;/code&gt;&lt;/strong&gt; folds the ledger to a current state and returns a boolean plus a reason. Every job that touches personal data calls it — ingestion, transformation, activation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deny-by-default.&lt;/strong&gt; No matching allow event means &lt;em&gt;not allowed&lt;/em&gt;. This is what makes DPDP's "consent required" and GDPR's "lawful basis required" the same code path, while CCPA's "allowed unless opted out" is expressed as a default-allow with an opt-out event.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvsdl5it85u43surw77a9.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvsdl5it85u43surw77a9.jpeg" alt="Iconographic unified consent registry diagram — three region identities resolving to one canonical subject_id, an append-only consent ledger with per-jurisdiction lawful-basis / opt-out / consent rows, and a withdrawal event flowing in." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — folding a consent ledger to a decision
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday operation is: given a subject, a purpose, and a jurisdiction, fold the event history into a yes/no. The fold encodes the &lt;em&gt;regime difference&lt;/em&gt; — GDPR/DPDP require a positive consent event, CCPA requires the absence of an opt-out event — so the same function serves all three.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Given the ledger below, may you run &lt;code&gt;marketing&lt;/code&gt; analytics for &lt;code&gt;s_88f1&lt;/code&gt; under GDPR, and under CCPA?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;subject_id&lt;/th&gt;
&lt;th&gt;jurisdiction&lt;/th&gt;
&lt;th&gt;purpose&lt;/th&gt;
&lt;th&gt;event&lt;/th&gt;
&lt;th&gt;ts&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;s_88f1&lt;/td&gt;
&lt;td&gt;GDPR&lt;/td&gt;
&lt;td&gt;marketing&lt;/td&gt;
&lt;td&gt;consent_granted&lt;/td&gt;
&lt;td&gt;2026-01-10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;s_88f1&lt;/td&gt;
&lt;td&gt;GDPR&lt;/td&gt;
&lt;td&gt;marketing&lt;/td&gt;
&lt;td&gt;consent_withdrawn&lt;/td&gt;
&lt;td&gt;2026-03-01&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;s_88f1&lt;/td&gt;
&lt;td&gt;CCPA&lt;/td&gt;
&lt;td&gt;sale&lt;/td&gt;
&lt;td&gt;opt_out&lt;/td&gt;
&lt;td&gt;2026-02-15&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;collections&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;defaultdict&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;is_allowed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;purpose&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;jurisdiction&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# fold history -&amp;gt; current state; deny-by-default for opt-in regimes
&lt;/span&gt;    &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subject_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;subject_id&lt;/span&gt;
            &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jurisdiction&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;jurisdiction&lt;/span&gt;
            &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;purpose&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;purpose&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sale&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
    &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;            &lt;span class="c1"&gt;# chronological fold
&lt;/span&gt;    &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;jurisdiction&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GDPR&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DPDP&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;consent_granted&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;       &lt;span class="c1"&gt;# positive consent required
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;jurisdiction&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CCPA&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;opt_out&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;               &lt;span class="c1"&gt;# allowed unless opted out
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; For &lt;strong&gt;GDPR&lt;/strong&gt; the fold sees &lt;code&gt;consent_granted&lt;/code&gt; then &lt;code&gt;consent_withdrawn&lt;/code&gt;; the last state is &lt;code&gt;consent_withdrawn&lt;/code&gt;, and since GDPR needs a positive &lt;code&gt;consent_granted&lt;/code&gt;, the answer is &lt;em&gt;no&lt;/em&gt;. For &lt;strong&gt;CCPA&lt;/strong&gt; the relevant &lt;code&gt;sale&lt;/code&gt; events fold to &lt;code&gt;opt_out&lt;/code&gt;, and because CCPA is allowed-unless-opted-out, &lt;code&gt;state == "opt_out"&lt;/code&gt; returns &lt;em&gt;no&lt;/em&gt; as well. The identical function returns the right answer for opposite legal models because the fold rule branches on jurisdiction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;jurisdiction&lt;/th&gt;
&lt;th&gt;folded state&lt;/th&gt;
&lt;th&gt;marketing allowed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GDPR&lt;/td&gt;
&lt;td&gt;consent_withdrawn&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CCPA&lt;/td&gt;
&lt;td&gt;opt_out&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Store consent as an immutable event stream and compute the boolean on read — a mutable &lt;code&gt;consent&lt;/code&gt; column throws away the history that every audit and every "was this legal at the time?" question depends on.&lt;/p&gt;

&lt;h3&gt;
  
  
  Privacy-engineering interview question on the consent registry
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; An interviewer gives you three source systems — an EU billing DB keyed by &lt;code&gt;eu_customer_id&lt;/code&gt;, a California app keyed by &lt;code&gt;ca_user_id&lt;/code&gt;, and an Indian app keyed by &lt;code&gt;in_principal_id&lt;/code&gt; — plus a verified-contact table. Design the resolution so a single erasure request deletes the person everywhere, and show the code that produces the canonical &lt;code&gt;subject_id&lt;/code&gt; and rejects an unverified match.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using deterministic identity resolution on verified contacts
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;canonical_subject_id&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;identity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;verified_contacts&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Only *verified* email/phone may bridge identities across regimes.
&lt;/span&gt;    &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;contact&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phone&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;val&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;identity&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;contact&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;val&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;verified_contacts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;contact&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;val&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;   &lt;span class="c1"&gt;# must be verified
&lt;/span&gt;            &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;contact&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;val&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# No verified bridge -&amp;gt; do NOT merge; mint a region-local surrogate.
&lt;/span&gt;        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;local:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;identity&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;identity&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s_&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()[:&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;source&lt;/th&gt;
&lt;th&gt;source_id&lt;/th&gt;
&lt;th&gt;verified contact&lt;/th&gt;
&lt;th&gt;resolves to&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;EU billing&lt;/td&gt;
&lt;td&gt;eu_1001&lt;/td&gt;
&lt;td&gt;email &lt;a href="mailto:priya@ex.com"&gt;priya@ex.com&lt;/a&gt; ✓&lt;/td&gt;
&lt;td&gt;s_88f1c2a4e9b0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CA app&lt;/td&gt;
&lt;td&gt;ca_7742&lt;/td&gt;
&lt;td&gt;email &lt;a href="mailto:priya@ex.com"&gt;priya@ex.com&lt;/a&gt; ✓&lt;/td&gt;
&lt;td&gt;s_88f1c2a4e9b0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IN app&lt;/td&gt;
&lt;td&gt;in_5590&lt;/td&gt;
&lt;td&gt;phone unverified ✗&lt;/td&gt;
&lt;td&gt;local:IN:in_5590&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Each incoming identity is looked up against the &lt;code&gt;verified_contacts&lt;/code&gt; set — a contact only bridges regimes if it was actually confirmed (double opt-in, OTP), never on a raw string match.&lt;/li&gt;
&lt;li&gt;The EU and CA rows both carry a &lt;em&gt;verified&lt;/em&gt; &lt;code&gt;priya@ex.com&lt;/code&gt;, so both hash to the &lt;strong&gt;same&lt;/strong&gt; &lt;code&gt;subject_id&lt;/code&gt;; a later erasure on that id reaches both systems.&lt;/li&gt;
&lt;li&gt;The IN row's phone is unverified, so it does &lt;strong&gt;not&lt;/strong&gt; merge — merging on an unverified identifier would risk deleting or exposing the wrong person, a worse failure than an unmerged duplicate.&lt;/li&gt;
&lt;li&gt;The hash is deterministic and one-way, so the surrogate is stable across runs yet never leaks the underlying email.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;canonical subject_id&lt;/th&gt;
&lt;th&gt;linked source records&lt;/th&gt;
&lt;th&gt;mergeable&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;s_88f1c2a4e9b0&lt;/td&gt;
&lt;td&gt;eu_1001, ca_7742&lt;/td&gt;
&lt;td&gt;yes (verified email)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;local:IN:in_5590&lt;/td&gt;
&lt;td&gt;in_5590&lt;/td&gt;
&lt;td&gt;no (unverified)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Deterministic resolution&lt;/strong&gt;&lt;/strong&gt; — hashing a verified, normalised contact yields the same &lt;code&gt;subject_id&lt;/code&gt; on every run, so identity is reproducible and joinable without central coordination.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Verify before you merge&lt;/strong&gt;&lt;/strong&gt; — bridging only on confirmed contacts prevents a mis-merge that would let one person's erasure or access request hit another's data — the highest-severity privacy bug.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Surrogate over PII&lt;/strong&gt;&lt;/strong&gt; — the platform joins on &lt;code&gt;subject_id&lt;/code&gt;, never on email, so analytics never touches a regulated attribute and the join key itself is not personal data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Fail-safe default&lt;/strong&gt;&lt;/strong&gt; — an unverified match stays a separate region-local id; a duplicate is recoverable, a wrong merge is a breach.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — resolution is O(1) per identity (a hash and a set lookup), so it scales linearly with ingest volume and adds no join-time overhead.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Access control&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — access-control&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Consent-gate and access-control problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Data quality&lt;/span&gt;
&lt;span&gt;Topic — data-quality&lt;/span&gt;
&lt;strong&gt;Identity-resolution and dedup problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-quality" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. Policy tags, data mapping &amp;amp; purpose limitation
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Tags on the schema turn compliance into a query, not a code review
&lt;/h3&gt;

&lt;p&gt;The move that scales privacy across hundreds of tables is &lt;strong&gt;policy-as-code tags attached to every column, so the pipeline enforces jurisdiction, purpose, and retention by reading the tag instead of by hand-written rules per table&lt;/strong&gt;. All three regimes demand a form of this: GDPR's Article 30 requires a &lt;strong&gt;Record of Processing Activities (RoPA)&lt;/strong&gt; — a living map of what data you hold, why, and for how long; DPDP requires an itemised notice inventory; CCPA requires you to catalogue categories of personal information and their purposes. The data map is that record made queryable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tag the columns, not the docs.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;pii_class&lt;/code&gt;.&lt;/strong&gt; Marks a column as &lt;code&gt;direct&lt;/code&gt; PII (email, phone, national id), &lt;code&gt;indirect&lt;/code&gt;/quasi-identifier (ip, device, zip), &lt;code&gt;sensitive&lt;/code&gt; (health, biometric, financial — GDPR Article 9 special category, CPRA "sensitive PI", DPDP has no separate sensitive tier but treats children's data specially), or &lt;code&gt;non_pii&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;jurisdiction&lt;/code&gt;.&lt;/strong&gt; Which regimes the column's rows fall under, derived from the subject registry, so a query can filter to EU-only rows when a rule is EU-specific.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;purpose_allowed&lt;/code&gt;.&lt;/strong&gt; The closed set of purposes this column may be used for — &lt;code&gt;billing&lt;/code&gt;, &lt;code&gt;fraud&lt;/code&gt;, &lt;code&gt;analytics&lt;/code&gt;, &lt;code&gt;marketing&lt;/code&gt;. This is the machine-readable form of &lt;strong&gt;purpose limitation&lt;/strong&gt; (GDPR Article 5(1)(b)): data collected for one purpose may not be silently repurposed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;retention&lt;/code&gt;.&lt;/strong&gt; A duration after which the value must be tombstoned — the operational form of GDPR's storage limitation and DPDP's "erase when the purpose is served".&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The data map is a catalog you can query.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Who / what / why / where / how long.&lt;/strong&gt; Each dataset entry answers the RoPA questions in structured fields, so "which tables hold EU sensitive data used for marketing?" is a &lt;code&gt;SELECT&lt;/code&gt;, not a week of interviews.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drift detection.&lt;/strong&gt; New columns without tags fail CI. An untagged PII column is a compliance hole, so the catalog is enforced at schema-change time, the same way a data contract is.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Purpose limitation is enforced at read time.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A purpose token on every job.&lt;/strong&gt; A job declares the purpose it runs for (&lt;code&gt;purpose="fraud"&lt;/code&gt;); the access layer intersects that with each column's &lt;code&gt;purpose_allowed&lt;/code&gt; and strips or blocks columns that do not match.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repurposing requires a new basis.&lt;/strong&gt; Using &lt;code&gt;marketing&lt;/code&gt;-tagged data for a &lt;code&gt;fraud&lt;/code&gt; model is not a bug you catch in review — the token mismatch blocks it, and lifting the block requires a new consent/basis record, which is exactly the legal reality.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuc83q2xksqrsyhgcc8g2.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuc83q2xksqrsyhgcc8g2.jpeg" alt="Iconographic policy-tag diagram — warehouse columns tagged with pii_class, jurisdiction, purpose_allowed and retention, a queryable data map catalog, and a purpose-token gate deciding which columns a job may read." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — a purpose-limited column projection
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The clearest way to see purpose limitation is a function that, given a job's purpose, returns only the columns that purpose is allowed to read. The tags live with the schema; the gate is a set intersection. This is how you stop a marketing job from quietly training on billing data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A &lt;code&gt;fraud&lt;/code&gt; job asks to read the &lt;code&gt;transactions&lt;/code&gt; table. Given the column tags below, which columns does it receive?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;column&lt;/th&gt;
&lt;th&gt;pii_class&lt;/th&gt;
&lt;th&gt;purpose_allowed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;amount&lt;/td&gt;
&lt;td&gt;non_pii&lt;/td&gt;
&lt;td&gt;billing, fraud, analytics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;card_token&lt;/td&gt;
&lt;td&gt;indirect&lt;/td&gt;
&lt;td&gt;fraud&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;email&lt;/td&gt;
&lt;td&gt;direct&lt;/td&gt;
&lt;td&gt;billing, marketing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;device_ip&lt;/td&gt;
&lt;td&gt;indirect&lt;/td&gt;
&lt;td&gt;fraud, analytics&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;project_for_purpose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;schema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;purpose&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Return only columns whose purpose_allowed set includes this job's purpose.
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;column&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;col&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;schema&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;purpose&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;purpose_allowed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;visible&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;project_for_purpose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;transactions_schema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;purpose&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fraud&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The gate walks the column tags and keeps a column only if &lt;code&gt;"fraud"&lt;/code&gt; is in its &lt;code&gt;purpose_allowed&lt;/code&gt; set. &lt;code&gt;amount&lt;/code&gt;, &lt;code&gt;card_token&lt;/code&gt;, and &lt;code&gt;device_ip&lt;/code&gt; all list &lt;code&gt;fraud&lt;/code&gt;, so they pass. &lt;code&gt;email&lt;/code&gt; lists only &lt;code&gt;billing&lt;/code&gt; and &lt;code&gt;marketing&lt;/code&gt;, so the fraud job never sees it — even though the column physically sits in the same table. Purpose limitation becomes a projection, applied uniformly to every job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;column&lt;/th&gt;
&lt;th&gt;visible to fraud job&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;amount&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;card_token&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;email&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;device_ip&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If the only thing standing between a marketing job and PII is a reviewer noticing the wrong &lt;code&gt;SELECT&lt;/code&gt;, purpose limitation is not enforced — put the rule on the column and make the platform apply it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Privacy-engineering interview question on retention and the data map
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; GDPR storage limitation and DPDP's "erase when the purpose is served" both demand that data not outlive its purpose. You have thousands of tables. How do you drive retention off the data map so that expired rows are tombstoned automatically, and prove it happened for an audit?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using retention tags to schedule tombstones
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timedelta&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;rows_to_tombstone&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;catalog&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;today&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Emit (table, column, cutoff) work items from the data map's retention tags.
&lt;/span&gt;    &lt;span class="n"&gt;work&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;catalog&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;                       &lt;span class="c1"&gt;# catalog == queryable RoPA
&lt;/span&gt;        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;col&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;columns&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="n"&gt;days&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;retention_days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;days&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;continue&lt;/span&gt;
            &lt;span class="n"&gt;cutoff&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;today&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nf"&gt;timedelta&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;days&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;days&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;work&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;table&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;table&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;column&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;column&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cutoff&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;cutoff&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;               &lt;span class="c1"&gt;# tombstone rows older than this
&lt;/span&gt;                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;basis&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;storage_limitation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;work&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;table&lt;/th&gt;
&lt;th&gt;column&lt;/th&gt;
&lt;th&gt;retention_days&lt;/th&gt;
&lt;th&gt;cutoff (today=2026-09-15)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;events&lt;/td&gt;
&lt;td&gt;device_ip&lt;/td&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;td&gt;2026-06-17&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;profiles&lt;/td&gt;
&lt;td&gt;email&lt;/td&gt;
&lt;td&gt;730&lt;/td&gt;
&lt;td&gt;2024-09-16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ledger&lt;/td&gt;
&lt;td&gt;amount&lt;/td&gt;
&lt;td&gt;(none)&lt;/td&gt;
&lt;td&gt;skipped&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The scheduler reads the &lt;strong&gt;data map&lt;/strong&gt;, not the tables — retention lives with the catalog entry, so one job covers every dataset.&lt;/li&gt;
&lt;li&gt;For each column carrying a &lt;code&gt;retention_days&lt;/code&gt; tag it computes a cutoff date; rows whose event time predates the cutoff are due for a tombstone.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;amount&lt;/code&gt; has no retention tag, so it is skipped — financial records often have a &lt;em&gt;legal-obligation&lt;/em&gt; basis to be retained, which the absence of a tag encodes.&lt;/li&gt;
&lt;li&gt;Each emitted work item names a &lt;code&gt;basis&lt;/code&gt; (&lt;code&gt;storage_limitation&lt;/code&gt;), so the resulting tombstone log is a self-describing audit record: what was erased, from where, and under which legal rationale.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;work items emitted&lt;/th&gt;
&lt;th&gt;audit fields per item&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2 (device_ip, email)&lt;/td&gt;
&lt;td&gt;table, column, cutoff, basis, run_ts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Data map as driver&lt;/strong&gt;&lt;/strong&gt; — putting &lt;code&gt;retention_days&lt;/code&gt; on the catalog means one scheduler enforces storage limitation across thousands of tables, instead of a cron per table that drifts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Tag absence is a decision&lt;/strong&gt;&lt;/strong&gt; — an untagged column is deliberately retained (legal obligation), so the model distinguishes "keep forever on purpose" from "nobody set a rule", which CI catches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Basis on every action&lt;/strong&gt;&lt;/strong&gt; — stamping each tombstone with its legal basis turns the erasure log into the evidence an auditor and a DPO ask for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Uniform across regimes&lt;/strong&gt;&lt;/strong&gt; — the same cutoff mechanism serves GDPR storage limitation and DPDP purpose-fulfilment because both reduce to "erase rows older than the tagged horizon".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — building the work list is O(columns) over the catalog, and the actual erase is O(expired rows) per table, decoupled from total table size via a time index.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Data quality&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — data-quality&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Data-contract and retention-validation problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-quality" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Access control&lt;/span&gt;
&lt;span&gt;Topic — access-control&lt;/span&gt;
&lt;strong&gt;Purpose-limited access and column-projection problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Subject-rights pipelines — access &amp;amp; erasure tombstones
&lt;/h2&gt;
&lt;h3&gt;
  
  
  One router, three SLAs — access fans out and assembles, erasure writes a tombstone that propagates
&lt;/h3&gt;

&lt;p&gt;Subject rights are where a privacy platform earns its keep, and the design that satisfies all three regimes is &lt;strong&gt;a single request router that branches into an access pipeline and an erasure pipeline, both driven by the data map and both leaving an audit trail&lt;/strong&gt;. The regimes differ in deadline and scope — GDPR's DSAR is one month, CCPA's requests are forty-five days, DPDP routes through the fiduciary's grievance mechanism — but the &lt;em&gt;engineering&lt;/em&gt; is identical: find every copy of the subject, then either export it or erase it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The router normalises the request.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One request type, per-regime SLA.&lt;/strong&gt; A request carries &lt;code&gt;(subject_id, kind, jurisdiction, received_ts)&lt;/code&gt;. The router stamps the deadline from the jurisdiction (30 / 45 / grievance-window days) and enqueues the work. Verification of the requester happens here — a "verifiable consumer request" under CCPA, identity confirmation under GDPR/DPDP.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deny-by-default on verification.&lt;/strong&gt; An unverified request is never executed; responding to an impersonator is itself a breach.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Access: fan out over the data map, assemble a package.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The data map is the index of where to look.&lt;/strong&gt; Because every dataset holding PII is catalogued and keyed by &lt;code&gt;subject_id&lt;/code&gt;, "find everything about this person" is a fan-out over the map, not tribal knowledge of which tables matter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assemble, redact, deliver.&lt;/strong&gt; Results are collected into an export bundle; other subjects' data that co-occurs (a shared thread, a joint transaction) is redacted. GDPR portability additionally requires a structured, machine-readable format — a shape DPDP does not mandate and CCPA satisfies with the "right to know".&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Erasure: tombstone, then propagate.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A tombstone is a durable erasure marker, not a &lt;code&gt;DELETE&lt;/code&gt;.&lt;/strong&gt; Writing &lt;code&gt;{subject_id, erased_at, scope}&lt;/code&gt; to a tombstone table means the fact of erasure survives even as it propagates asynchronously to warehouses, lakes, backups, and search indices. A raw &lt;code&gt;DELETE&lt;/code&gt; on the primary leaves stale copies downstream that no one tracks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Propagation to every copy.&lt;/strong&gt; The tombstone is a change-data event every downstream consumer subscribes to; each applies the erasure to its own store and acknowledges. Erasure is complete only when every registered sink has acked — which is why the data map must also list sinks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Idempotent and re-runnable.&lt;/strong&gt; Applying the same tombstone twice is a no-op, so a retried or replayed erasure is safe; this is the property that lets you re-drive a failed propagation without fear.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhvh62rw5vik41v37levq.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhvh62rw5vik41v37levq.jpeg" alt="Iconographic subject-rights pipeline diagram — one request router branching to access (fan-out and assemble a package) and erasure (write a tombstone and propagate it to every downstream copy), with per-regime SLA clocks." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — an erasure tombstone with downstream propagation
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The canonical erasure is not a delete statement; it is an event that fans out. You write one tombstone, and every downstream store consumes it and confirms. The pipeline is done when all sinks ack, and the tombstone table is the proof.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A verified GDPR erasure arrives for &lt;code&gt;s_88f1&lt;/code&gt;. Show how the tombstone drives deletion across the primary DB, the warehouse, and a backup, and how you know it is complete.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;sink&lt;/th&gt;
&lt;th&gt;keyed by subject_id&lt;/th&gt;
&lt;th&gt;ack status before&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;primary_db&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;warehouse&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;backup&lt;/td&gt;
&lt;td&gt;yes (restore-time apply)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;erase&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sinks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tombstones&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# 1) Durable marker first, 2) fan out, 3) require all acks (idempotent).
&lt;/span&gt;    &lt;span class="n"&gt;tombstones&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subject_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                       &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;erased_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scope&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acks&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()})&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;sink&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;sinks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;sink&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;apply_tombstone&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;        &lt;span class="c1"&gt;# each is a no-op if already applied
&lt;/span&gt;        &lt;span class="n"&gt;tombstones&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acks&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sink&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;required&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;sinks&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;complete&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;required&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;issubset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tombstones&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acks&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;complete&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;action&lt;/th&gt;
&lt;th&gt;sink&lt;/th&gt;
&lt;th&gt;acks so far&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;write tombstone&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;{}&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;apply + ack&lt;/td&gt;
&lt;td&gt;primary_db&lt;/td&gt;
&lt;td&gt;{primary_db}&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;apply + ack&lt;/td&gt;
&lt;td&gt;warehouse&lt;/td&gt;
&lt;td&gt;{primary_db, warehouse}&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;apply + ack&lt;/td&gt;
&lt;td&gt;backup&lt;/td&gt;
&lt;td&gt;{primary_db, warehouse, backup}&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The tombstone is written &lt;strong&gt;before&lt;/strong&gt; any deletion, so if the process crashes mid-propagation the erasure is not lost — a restart re-drives it from the durable marker.&lt;/li&gt;
&lt;li&gt;Each sink applies the erasure and acks; because &lt;code&gt;apply_tombstone&lt;/code&gt; is idempotent, a re-driven step that already ran simply re-acks without error.&lt;/li&gt;
&lt;li&gt;Completion is &lt;code&gt;required.issubset(acks)&lt;/code&gt; — every registered sink must confirm, including the backup, which applies the tombstone at restore time so a recovered snapshot never resurrects the person.&lt;/li&gt;
&lt;li&gt;The tombstone row, with its ack set and timestamp, is the audit artifact proving the erasure met the regime's deadline.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;subject_id&lt;/th&gt;
&lt;th&gt;sinks acked&lt;/th&gt;
&lt;th&gt;complete&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;s_88f1&lt;/td&gt;
&lt;td&gt;primary_db, warehouse, backup&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Erasure is finished when the last downstream copy acks the tombstone, not when the primary &lt;code&gt;DELETE&lt;/code&gt; returns — model the sinks explicitly or you will leave the person in a lake or a backup.&lt;/p&gt;

&lt;h3&gt;
  
  
  Privacy-engineering interview question on cross-regime request routing
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Your inbox receives access and deletion requests for the same &lt;code&gt;subject_id&lt;/code&gt; across GDPR, CCPA, and DPDP, each with its own deadline. Design a router that assigns the correct SLA, refuses unverified requests, and lets an erasure be retried safely if a downstream sink was offline.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using an idempotent router keyed on request id
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;SLA_DAYS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GDPR&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CCPA&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;45&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DPDP&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;verified&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;done_ids&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Idempotent: a replayed request_id is skipped; unverified is refused.
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;request_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;done_ids&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;already_processed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;verified&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;request_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refused&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reason&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unverified requester&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;deadline&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;add_days&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;received_ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;SLA_DAYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jurisdiction&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]])&lt;/span&gt;
    &lt;span class="n"&gt;handler&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;access_pipeline&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kind&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;access&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="n"&gt;erasure_pipeline&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;handler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subject_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;jurisdiction&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jurisdiction&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;done_ids&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;request_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;         &lt;span class="c1"&gt;# mark handled -&amp;gt; safe to retry
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;completed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deadline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;deadline&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;detail&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;request_id&lt;/th&gt;
&lt;th&gt;jurisdiction&lt;/th&gt;
&lt;th&gt;kind&lt;/th&gt;
&lt;th&gt;verified&lt;/th&gt;
&lt;th&gt;outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;r1&lt;/td&gt;
&lt;td&gt;GDPR&lt;/td&gt;
&lt;td&gt;erasure&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;completed, deadline +30d&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;r1&lt;/td&gt;
&lt;td&gt;GDPR&lt;/td&gt;
&lt;td&gt;erasure&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;already_processed (replay)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;r2&lt;/td&gt;
&lt;td&gt;CCPA&lt;/td&gt;
&lt;td&gt;access&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;refused (unverified)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;r3&lt;/td&gt;
&lt;td&gt;DPDP&lt;/td&gt;
&lt;td&gt;erasure&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;completed, deadline +30d&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;r1&lt;/code&gt; is verified and new, so the router stamps a 30-day GDPR deadline and calls the erasure pipeline, then records &lt;code&gt;r1&lt;/code&gt; as done.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;replay&lt;/strong&gt; of &lt;code&gt;r1&lt;/code&gt; (a retried queue message) short-circuits on the &lt;code&gt;done_ids&lt;/code&gt; check, so the erasure is never applied twice — the router itself is idempotent even before the sinks are.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;r2&lt;/code&gt; is a CCPA access request but the requester failed verification, so it is refused; responding would have leaked data to an impersonator.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;r3&lt;/code&gt; is a DPDP erasure, verified, and gets the DPDP window — the same router serves three regimes by looking the SLA up from a table.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;request_id&lt;/th&gt;
&lt;th&gt;status&lt;/th&gt;
&lt;th&gt;deadline&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;r1&lt;/td&gt;
&lt;td&gt;completed / then already_processed&lt;/td&gt;
&lt;td&gt;+30d&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;r2&lt;/td&gt;
&lt;td&gt;refused&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;r3&lt;/td&gt;
&lt;td&gt;completed&lt;/td&gt;
&lt;td&gt;+30d&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Single router, table-driven SLA&lt;/strong&gt;&lt;/strong&gt; — the deadline is a lookup, so adding a fourth regime is a config row, not a new pipeline; the branch on &lt;code&gt;kind&lt;/code&gt; chooses access vs erasure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Idempotent on request_id&lt;/strong&gt;&lt;/strong&gt; — a &lt;code&gt;done_ids&lt;/code&gt; guard makes replays and retries no-ops, so an at-least-once queue never double-erases or double-exports.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Verify-then-act&lt;/strong&gt;&lt;/strong&gt; — refusing unverified requests is the router's most important job; a fast, wrong response to an impersonator is worse than a slow, correct one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Deadline as a first-class field&lt;/strong&gt;&lt;/strong&gt; — stamping the SLA on completion produces the timeliness evidence each regulator expects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — routing is O(1) per request; the real work is the O(copies) fan-out inside the handler, bounded by the data map's sink list.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Access control&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — access-control&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Subject-request routing and verification problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Tokenization&lt;/span&gt;
&lt;span&gt;Topic — tokenization&lt;/span&gt;
&lt;strong&gt;Tombstone-propagation and erasure problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/tokenization" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Tokenization &amp;amp; pseudonymization
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Keep raw PII in a vault, land tokens in the warehouse, and make erasure a single key delete
&lt;/h3&gt;

&lt;p&gt;The technique that makes every previous section cheaper is &lt;strong&gt;de-identification: replace raw PII with tokens so the warehouse holds no direct identifiers, and use per-subject keys so erasure becomes deleting one key rather than chasing rows across every store&lt;/strong&gt;. GDPR explicitly rewards &lt;strong&gt;pseudonymization&lt;/strong&gt; (Article 4(5) and Article 32) as a security measure; CPRA and DPDP both treat de-identified data as lower-risk. Getting the vocabulary exact is itself an interview filter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three terms interviewers make you separate.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tokenization — reversible, vault-backed.&lt;/strong&gt; The real value (a card number, an email) is stored in a secured vault and replaced everywhere else by an opaque &lt;code&gt;token&lt;/code&gt;. Given the vault and authorisation, you can re-identify. The warehouse never holds the raw value.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pseudonymization — reversible only with separate info (GDPR Art 4(5)).&lt;/strong&gt; The data can be attributed to a person &lt;em&gt;only&lt;/em&gt; by using additional information kept separately and under controls. Tokenization is one implementation; the legal point is that the mapping is held apart and protected. Pseudonymised data is still personal data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anonymization — irreversible.&lt;/strong&gt; No key, no vault, no path back to the person. Truly anonymised data falls outside GDPR/DPDP entirely — but true anonymisation is hard, because quasi-identifiers can re-identify, so most "anonymised" data is really pseudonymised.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Crypto-shredding makes erasure O(1).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A key per subject.&lt;/strong&gt; Each subject's PII is encrypted with a subject-specific key; the ciphertext can live anywhere, even in backups. To erase the person you &lt;strong&gt;delete their key&lt;/strong&gt;, and every ciphertext copy — warehouse, lake, backup, index — becomes permanently unreadable at once.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why this beats chasing rows.&lt;/strong&gt; The propagation problem from Section 4 shrinks: you still tombstone the identity, but the &lt;em&gt;content&lt;/em&gt; is neutralised the instant the key dies, so a missed backup copy is inert rather than a breach.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Deterministic vs randomized tokens — the analytics trade-off.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic tokens join.&lt;/strong&gt; The same input always maps to the same token, so you can &lt;code&gt;GROUP BY token&lt;/code&gt; and join across tables without re-identifying. The cost is that deterministic tokens are vulnerable to frequency analysis, so they suit low-cardinality-safe joins, not secrets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Randomized tokens do not join.&lt;/strong&gt; A fresh token per occurrence is safest but breaks joins, so it fits values you never aggregate on (a free-text note, a one-off id). Choosing per column is a data-modelling decision the tag from Section 3 can carry.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv44dz7t03c6hzj9ugpab.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv44dz7t03c6hzj9ugpab.jpeg" alt="Iconographic tokenization diagram — raw PII kept in a vault, tokens flowing to the warehouse, a per-subject key enabling crypto-shred erasure, and a purpose-gated re-identify path." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — tokenizing on ingest and re-identifying by purpose
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday pattern is: on ingest, swap each PII field for a deterministic token and stash the real value in the vault; downstream jobs work on tokens; a re-identify call, gated by purpose, is the only path back. This keeps the warehouse free of raw PII while preserving joins.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Tokenize an incoming record's &lt;code&gt;email&lt;/code&gt;, load only the token to the warehouse, and show what a &lt;code&gt;fraud&lt;/code&gt;-purpose re-identify returns versus a &lt;code&gt;marketing&lt;/code&gt; one when email is not marketing-allowed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"subject_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"s_88f1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"email"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"priya@example.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"amount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;42.5&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;hmac&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;

&lt;span class="n"&gt;VAULT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;  &lt;span class="c1"&gt;# token -&amp;gt; raw value, secured &amp;amp; access-controlled in reality
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;tokenize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subject-key-s_88f1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;tok&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tok_&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;hmac&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()[:&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;VAULT&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;tok&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;                          &lt;span class="c1"&gt;# raw stays in the vault only
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;tok&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;reidentify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;purpose&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;allowed_purposes&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;purpose&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;allowed_purposes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()):&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;PermissionError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;purpose not permitted for this field&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;VAULT&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;row&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subject_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s_88f1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;tokenize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;priya@example.com&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;42.5&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;tokenize&lt;/code&gt; computes a deterministic HMAC token and files the raw email in the vault; the row written to the warehouse carries &lt;code&gt;email = tok_...&lt;/code&gt;, never the address. A &lt;code&gt;fraud&lt;/code&gt; job that is permitted to re-identify the email gets the raw value back; a &lt;code&gt;marketing&lt;/code&gt; job, for which email is not an allowed purpose, hits the &lt;code&gt;PermissionError&lt;/code&gt; and only ever sees the token. Erasure later deletes &lt;code&gt;subject-key-s_88f1&lt;/code&gt;, after which the token can no longer be regenerated or matched.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;caller purpose&lt;/th&gt;
&lt;th&gt;re-identify email&lt;/th&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;fraud&lt;/td&gt;
&lt;td&gt;allowed&lt;/td&gt;
&lt;td&gt;&lt;code&gt;priya@example.com&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;marketing&lt;/td&gt;
&lt;td&gt;not allowed&lt;/td&gt;
&lt;td&gt;PermissionError (token only)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If raw PII reaches the warehouse "just in case", you have inverted the model — tokenize at the edge and make re-identification the rare, purpose-gated exception.&lt;/p&gt;

&lt;h3&gt;
  
  
  Privacy-engineering interview question on erasure at scale
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You hold a subject's data across a warehouse, a data lake, and nightly backups. A GDPR erasure must guarantee the data is unrecoverable, but rewriting every backup is infeasible. How do you make erasure both complete and cheap?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using crypto-shredding with per-subject keys
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;KEYS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;   &lt;span class="c1"&gt;# subject_id -&amp;gt; encryption key (the only thing that must truly die)
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;store_pii&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setdefault&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;generate_key&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;encrypt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                  &lt;span class="c1"&gt;# ciphertext may live anywhere
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;crypto_shred&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Erasure = destroy the key; all ciphertext copies become unreadable.
&lt;/span&gt;    &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;read_pii&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ciphertext&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;                             &lt;span class="c1"&gt;# shredded -&amp;gt; permanently opaque
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;decrypt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ciphertext&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;action&lt;/th&gt;
&lt;th&gt;key present&lt;/th&gt;
&lt;th&gt;read result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;store_pii(s_88f1, email)&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;ciphertext in warehouse, lake, backup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;read_pii before shred&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;&lt;code&gt;priya@example.com&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;crypto_shred(s_88f1)&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;read_pii after shred (any copy)&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;None (unrecoverable)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Each subject's PII is encrypted with a &lt;strong&gt;per-subject key&lt;/strong&gt;, so the ciphertext is safe to replicate to the lake and to backups.&lt;/li&gt;
&lt;li&gt;Before erasure, an authorised read decrypts normally because the key exists.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;crypto_shred&lt;/code&gt; deletes only the key — a single, cheap operation — instead of hunting and rewriting every ciphertext copy.&lt;/li&gt;
&lt;li&gt;After shredding, &lt;em&gt;every&lt;/em&gt; copy, including untouched backups, decrypts to &lt;code&gt;None&lt;/code&gt;; the data is unrecoverable without ever rewriting a backup file.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;after crypto_shred&lt;/th&gt;
&lt;th&gt;warehouse&lt;/th&gt;
&lt;th&gt;lake&lt;/th&gt;
&lt;th&gt;backup&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;decryptable&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Key as the unit of erasure&lt;/strong&gt;&lt;/strong&gt; — destroying one per-subject key neutralises unbounded ciphertext copies at once, so erasure cost is O(1) rather than O(copies).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Backups without a rewrite&lt;/strong&gt;&lt;/strong&gt; — because a stale backup holds only ciphertext, a shredded key makes that backup inert without touching it — the only practical way to honour erasure against immutable backups.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Pseudonymization by construction&lt;/strong&gt;&lt;/strong&gt; — the warehouse holds ciphertext/tokens, so it is pseudonymised data under GDPR Art 4(5), reducing breach exposure even before any request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Tombstone still needed for identity&lt;/strong&gt;&lt;/strong&gt; — crypto-shred kills the content; you still write the Section 4 tombstone so the &lt;em&gt;fact&lt;/em&gt; of erasure and the identity linkage are removed and audited.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — shredding is O(1) (one key delete); the only ongoing cost is per-read decryption, an acceptable trade for making erasure and replication cheap.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Tokenization&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — tokenization&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Tokenization, vaulting and crypto-shred problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/tokenization" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Data quality&lt;/span&gt;
&lt;span&gt;Topic — data-quality&lt;/span&gt;
&lt;strong&gt;De-identification and re-identification-risk problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-quality" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — multi-jurisdiction privacy recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Regime field mapping.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;REGIME&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GDPR&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subject&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data subject&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;party&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;controller&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trigger&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lawful basis (Art 6)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sale_optout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;erasure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;right to be forgotten (Art 17)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;access_sla&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CCPA&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subject&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;consumer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;party&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;business&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trigger&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notice at collection&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sale_optout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;erasure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;right to delete (exceptions)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;access_sla&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;45&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DPDP&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subject&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data principal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;party&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data fiduciary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trigger&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;consent / legitimate use&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sale_optout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;erasure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;on withdrawal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;access_sla&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Consent-check gate (deny-by-default for opt-in regimes).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;allowed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;jurisdiction&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;jurisdiction&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GDPR&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DPDP&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;consent_granted&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;jurisdiction&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CCPA&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;opt_out&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Policy-tag column contract.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;column&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pii_class&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;direct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jurisdiction&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GDPR&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CCPA&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DPDP&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
          &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;purpose_allowed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;billing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;retention_days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;730&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Erasure tombstone.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;tombstone&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subject_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s_88f1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;erased_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scope&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acks&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()}&lt;/span&gt;   &lt;span class="c1"&gt;# complete when all sinks ack
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Crypto-shred (O(1) erasure).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;crypto_shred&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="c1"&gt;# every ciphertext copy goes dark at once
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Purpose-limited read.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;visible&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;schema&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;purpose&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;purpose_allowed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Regime comparison at a glance.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;dimension&lt;/th&gt;
&lt;th&gt;GDPR (EU)&lt;/th&gt;
&lt;th&gt;CCPA/CPRA (California)&lt;/th&gt;
&lt;th&gt;DPDP (India)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;individual&lt;/td&gt;
&lt;td&gt;data subject&lt;/td&gt;
&lt;td&gt;consumer&lt;/td&gt;
&lt;td&gt;data principal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;processing trigger&lt;/td&gt;
&lt;td&gt;lawful basis&lt;/td&gt;
&lt;td&gt;notice + opt-out&lt;/td&gt;
&lt;td&gt;consent / legitimate use&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;access deadline&lt;/td&gt;
&lt;td&gt;1 month&lt;/td&gt;
&lt;td&gt;45 days&lt;/td&gt;
&lt;td&gt;rules-set window&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;erasure&lt;/td&gt;
&lt;td&gt;broad (Art 17)&lt;/td&gt;
&lt;td&gt;delete with exceptions&lt;/td&gt;
&lt;td&gt;on withdrawal&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is a multi-jurisdiction privacy pipeline?
&lt;/h3&gt;

&lt;p&gt;It is a data platform designed so the same personal data can be processed legally under several privacy laws at once — typically EU GDPR, California CCPA/CPRA, and India's DPDP Act 2023 — without forking the infrastructure per region. The differences between the laws are pushed into data: a consent ledger, per-column policy tags, jurisdiction fields, and subject-rights pipelines that read those structures at runtime. The goal is one physical platform whose behaviour is parameterised by jurisdiction rather than hard-coded to one statute.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do GDPR, CCPA, and DPDP differ for a data engineer?
&lt;/h3&gt;

&lt;p&gt;The biggest engineering difference is the legal trigger for processing. GDPR requires a lawful basis (often opt-in consent) before you process; CCPA requires no consent but a notice at collection plus an opt-out of sale or sharing; DPDP requires consent (or a listed legitimate use) plus an itemised notice. They also differ on vocabulary (data subject vs consumer vs data principal), response deadlines (1 month vs 45 days vs a rules-set window), and rights (GDPR has portability, DPDP does not; CCPA has no lawful-basis concept). A design that assumes opt-in everywhere is wrong for California, and one that assumes opt-out is wrong for the EU and India.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between tokenization and pseudonymization?
&lt;/h3&gt;

&lt;p&gt;Tokenization is a concrete technique: replace a raw value with an opaque token and keep the real value in a secured vault, so the token is reversible only through that vault. Pseudonymization is the legal concept (GDPR Article 4(5)): data that can be attributed to a person only with additional information kept separately and protected — tokenization is one way to achieve it. Both are reversible and both leave the data as personal data; only true anonymization, which removes any path back to the individual, takes data outside GDPR and DPDP.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you engineer the right to erasure across downstream copies?
&lt;/h3&gt;

&lt;p&gt;Do not model erasure as a &lt;code&gt;DELETE&lt;/code&gt; on one table. Write a durable tombstone marker for the subject, then propagate it as an event to every downstream sink — warehouse, lake, search index, backups — and treat the erasure as complete only when every registered sink acknowledges. To handle backups you cannot rewrite, combine the tombstone with crypto-shredding: encrypt each subject's data with a per-subject key and delete the key, which renders every ciphertext copy, including old backups, permanently unreadable in one cheap operation.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is purpose limitation and how do you enforce it in a pipeline?
&lt;/h3&gt;

&lt;p&gt;Purpose limitation (GDPR Article 5(1)(b), echoed by DPDP) means data collected for one purpose may not be silently reused for another. You enforce it by tagging each column with the closed set of purposes it may serve, and by having every job declare the purpose it runs for; the access layer intersects the two and strips or blocks columns whose purpose does not match. Repurposing then requires a new consent or basis record rather than a code change a reviewer might miss, which mirrors the legal reality.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I need separate warehouses per jurisdiction?
&lt;/h3&gt;

&lt;p&gt;Usually no, and forking per region is the anti-pattern this whole design avoids. A single platform with a canonical &lt;code&gt;subject_id&lt;/code&gt;, a consent ledger, jurisdiction tags, and rights pipelines handles all three regimes and stays maintainable. You may still need data-residency controls — DPDP can restrict transfers to blacklisted countries and some GDPR flows require EU storage — but residency is a placement policy on tagged data, not a reason to duplicate the entire stack and its logic three times.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every idea above, from the consent-ledger fold and the purpose-limited projection to the erasure tombstone and crypto-shred, maps to a hands-on practice room where you build the pipeline against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you make this erasure complete across every copy?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;Practice access-control problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/tokenization" rel="noopener noreferrer"&gt;Tokenization drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>SOC 2 for Data Engineering: Controls, Audit Logs &amp; Evidence-Collection Pipelines</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Fri, 25 Sep 2026 13:37:58 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/soc-2-for-data-engineering-controls-audit-logs-evidence-collection-pipelines-5m8</link>
      <guid>https://dev.to/gowthampotureddi/soc-2-for-data-engineering-controls-audit-logs-evidence-collection-pipelines-5m8</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;soc 2 for data engineering&lt;/code&gt;&lt;/strong&gt; is the moment an abstract compliance framework turns into concrete work on your warehouse: who can read the PII column, how a terminated employee's grants get revoked, whether last Tuesday's pipeline deploy went through review, and how you prove all of it to an auditor without a week of screenshots. SOC 2 is not a product you install and not a certification you pass once — it is an independent attestation, written by a licensed CPA firm, that the controls you claim to run actually exist and actually operated over a window of time. For a data engineer, that window is measured against tables, grants, and query logs you already own.&lt;/p&gt;

&lt;p&gt;The confusion that trips people up is treating SOC 2 as a security-team problem that lands in a slide deck. It is not. The controls that get sampled hardest — least-privilege access to the warehouse, an audit trail of every read against sensitive data, change management on the pipelines that move that data, and evidence that survives tampering — are all data-engineering surfaces. This guide walks through the four things an auditor (and an interviewer) will actually probe: the access controls you provision and deprovision, the warehouse audit logs that prove them, the change-management chain behind every pipeline deploy, and the evidence-collection pipelines that gather it all automatically. Each section pairs the concept with a Solution-Tail interview answer — code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj8vxg00bgbwcy289n852.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj8vxg00bgbwcy289n852.jpeg" alt="PipeCode blog header for SOC 2 for data engineering — bold white headline 'SOC 2 for Data Engineering' with subtitle 'Controls · Audit Logs · Evidence Pipelines' and a stylised control-to-evidence scene on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;access-control practice library →&lt;/a&gt;, rehearse the monitoring and alerting side on the &lt;a href="https://pipecode.ai/explore/practice/topic/sla-monitoring" rel="noopener noreferrer"&gt;SLA-monitoring practice set →&lt;/a&gt;, and harden the correctness of your evidence tables on the &lt;a href="https://pipecode.ai/explore/practice/topic/data-quality" rel="noopener noreferrer"&gt;data-quality practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why SOC 2 lands on the data engineer's desk in 2026&lt;/li&gt;
&lt;li&gt;Access controls data engineers own&lt;/li&gt;
&lt;li&gt;Warehouse audit logs &amp;amp; access history&lt;/li&gt;
&lt;li&gt;Change management for pipelines&lt;/li&gt;
&lt;li&gt;Evidence-collection pipelines&lt;/li&gt;
&lt;li&gt;Cheat sheet — SOC 2 evidence recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why SOC 2 lands on the data engineer's desk in 2026
&lt;/h2&gt;

&lt;h3&gt;
  
  
  SOC 2 is an attestation against the Trust Services Criteria — not a checklist you buy, and half the sampled controls live in your warehouse
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;SOC 2 is a CPA firm's opinion that the controls you described are suitably designed and — for a Type II — actually operated over a period, and the data platform is where most of those controls physically live&lt;/strong&gt;. There is no "SOC 2 software" that makes you compliant; there is a set of controls you run, and an audit that samples evidence they ran. Understanding that shape is what separates a data engineer who can speak to auditors from one who forwards every question to the security team.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What SOC 2 actually is.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;An attestation, not a certification.&lt;/strong&gt; The output is a report signed by a licensed CPA firm following AICPA attestation standards, expressing an opinion on your controls. You do not "get certified"; you receive a report you can share with customers under NDA.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scoped to a system.&lt;/strong&gt; The report covers a defined system — for a data platform that usually means the warehouse, the pipelines, the orchestration, the access model, and the cloud infrastructure underneath. Anything out of scope is not covered.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Driven by your own control descriptions.&lt;/strong&gt; You write the controls (the "system description"); the auditor tests them. That is why vague controls hurt you — you will be tested against exactly what you claimed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The five Trust Services Criteria (TSC).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Security (the Common Criteria) — mandatory.&lt;/strong&gt; Every SOC 2 includes it: logical access, change management, risk assessment, monitoring. This is where most data-engineering controls sit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Availability — optional.&lt;/strong&gt; Uptime, backup, disaster recovery, capacity — relevant if you promise SLAs on data delivery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Processing Integrity — optional.&lt;/strong&gt; Data is processed completely, accurately, and on time — a natural fit for pipeline correctness and data-quality checks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confidentiality — optional.&lt;/strong&gt; Information designated confidential is protected — encryption, access restriction, retention/disposal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privacy — optional.&lt;/strong&gt; Personal information is collected, used, retained, and disposed of per your notice — the strictest, and often deferred.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Type I vs Type II — the distinction interviewers love.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Type I is a point in time.&lt;/strong&gt; The auditor opines that controls are &lt;em&gt;suitably designed&lt;/em&gt; as of a single date. It answers "do the right controls exist?" — a snapshot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Type II is a period.&lt;/strong&gt; The auditor opines that controls &lt;em&gt;operated effectively&lt;/em&gt; across a window, usually 3 to 12 months. It answers "did the controls actually run, every time, the whole period?" — and it is tested by sampling evidence across that window.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why it matters to a DE.&lt;/strong&gt; Type II is the one customers ask for, and it is the one that forces you to &lt;em&gt;retain&lt;/em&gt; evidence continuously. A control that worked once but has no audit trail across the period fails a Type II even if it was perfectly designed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The controls a data engineer actually owns.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Logical access&lt;/strong&gt; — provisioning, least privilege, periodic access reviews, deprovisioning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit logging&lt;/strong&gt; — warehouse query and access history proving who read and changed what.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Change management&lt;/strong&gt; — version control, peer review, and approvals on pipeline code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Encryption&lt;/strong&gt; — at rest and in transit, usually inherited from the cloud warehouse but attested by you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitoring&lt;/strong&gt; — pipeline failure alerting, anomaly detection, and control-monitoring dashboards.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evidence&lt;/strong&gt; — the pipelines that gather all of the above into a form an auditor can sample.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you say &lt;strong&gt;"SOC 2 is an attestation against the Trust Services Criteria,"&lt;/strong&gt; not "a certification you pass"? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you name &lt;strong&gt;Security as the mandatory Common Criteria&lt;/strong&gt; and the other four as scope-dependent? — required framing.&lt;/li&gt;
&lt;li&gt;Do you distinguish &lt;strong&gt;Type II as operating-effectiveness over a period&lt;/strong&gt; from Type I's point-in-time design? — the classic discriminator.&lt;/li&gt;
&lt;li&gt;Do you frame your job as &lt;strong&gt;"owning the access, logging, change-management, and evidence controls,"&lt;/strong&gt; not "helping the security team"? — ownership signal.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — mapping one control to its evidence
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The single most useful habit for SOC 2 is to stop thinking in vague policies and start thinking in &lt;em&gt;control → evidence&lt;/em&gt; pairs. Every control you claim must produce an artifact an auditor can pull. Modeling that mapping as data — a small control catalog — is the mindset that makes the rest of this guide click, because each later section is just one row of this catalog turned into a real query.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Express three data-platform controls as a catalog that pairs each control with the concrete evidence query or artifact that proves it operated.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;control_id&lt;/th&gt;
&lt;th&gt;control&lt;/th&gt;
&lt;th&gt;trust criterion&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CC6.1&lt;/td&gt;
&lt;td&gt;Least-privilege access to the warehouse&lt;/td&gt;
&lt;td&gt;Security&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CC6.2&lt;/td&gt;
&lt;td&gt;Access removed on termination&lt;/td&gt;
&lt;td&gt;Security&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CC8.1&lt;/td&gt;
&lt;td&gt;Changes are peer-reviewed and approved&lt;/td&gt;
&lt;td&gt;Security&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;control_catalog&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;   &lt;span class="c1"&gt;# each control maps to the evidence that proves it ran
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CC6.1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Least-privilege access to the warehouse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;evidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT grantee_name, role, privilege FROM account_usage.grants_to_users&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cadence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;quarterly access review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CC6.2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Access removed on termination&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;evidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;grants anti-joined against the active-employee roster (0 orphans)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cadence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;daily deprovisioning check&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CC8.1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Changes are peer-reviewed and approved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;evidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deployments LEFT JOIN approved_prs (0 unapproved prod deploys)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cadence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;per deploy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;cid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;control_catalog&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;evidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Each key is a control identifier borrowed from the Security Common Criteria numbering (&lt;code&gt;CC6.x&lt;/code&gt; is logical access, &lt;code&gt;CC8.x&lt;/code&gt; is change management). Each value pins the &lt;em&gt;evidence&lt;/em&gt; — an actual query or artifact — and a &lt;em&gt;cadence&lt;/em&gt; that says how often the evidence is produced. The auditor does not accept "we use least privilege"; they accept the output of the grants query, sampled on the review date. Turning every control into a runnable evidence statement is exactly what an evidence-collection pipeline automates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;control_id&lt;/th&gt;
&lt;th&gt;evidence the auditor samples&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CC6.1&lt;/td&gt;
&lt;td&gt;grant listing from &lt;code&gt;account_usage.grants_to_users&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CC6.2&lt;/td&gt;
&lt;td&gt;anti-join showing 0 orphaned grants&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CC8.1&lt;/td&gt;
&lt;td&gt;join showing every prod deploy has an approved PR&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If you cannot name the query or artifact that proves a control, you do not have a SOC 2 control — you have a policy. Every control must resolve to evidence you can pull on demand.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Access controls data engineers own
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Provisioning, least privilege, and deprovisioning are the CC6 controls auditors sample first — and the warehouse is where they live
&lt;/h3&gt;

&lt;p&gt;The Security criterion's logical-access controls (the &lt;code&gt;CC6&lt;/code&gt; family) are the ones every SOC 2 tests, and for a data team they resolve to warehouse grants. Say the loop in one breath: &lt;strong&gt;access is provisioned from an authoritative source, granted at least privilege through roles, reviewed periodically, and revoked the moment someone leaves&lt;/strong&gt;. Break any link and you have a finding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Provisioning from an authoritative source.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Joiner-mover-leaver.&lt;/strong&gt; Access should be driven by the HR system or an identity provider, not by ad-hoc Slack requests. A joiner gets a role, a mover's role changes, a leaver's access is revoked — each event tied to a record.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Request and approval trail.&lt;/strong&gt; Every grant should trace to an approved request (a ticket, an IdP group assignment) so the auditor can see &lt;em&gt;why&lt;/em&gt; someone has access, not just &lt;em&gt;that&lt;/em&gt; they do.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Least privilege through roles, not direct grants.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RBAC, never user grants.&lt;/strong&gt; In Snowflake, BigQuery, or Redshift, privileges attach to roles and roles attach to users. Granting &lt;code&gt;SELECT&lt;/code&gt; directly to a user is un-auditable and un-revocable at scale; granting a role is both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Separate read, write, and admin.&lt;/strong&gt; A &lt;code&gt;read_raw&lt;/code&gt; role, a &lt;code&gt;read_pii&lt;/code&gt; role gated behind a masking policy, and a &lt;code&gt;loader&lt;/code&gt; role that can write are three different privileges. Nobody gets the union "just in case."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Masking and row-access policies.&lt;/strong&gt; Column masking and row-access policies let you grant a role table access while still restricting the sensitive columns or rows — least privilege at the cell level, and itself an attestable control.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Periodic access reviews.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Quarterly is the common cadence.&lt;/strong&gt; A reviewer confirms each grant is still needed. The evidence is the review record plus the grant snapshot it reviewed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The review must be actionable.&lt;/strong&gt; "Looks fine" is not evidence; a diff showing what changed since last quarter, and any revocations that resulted, is.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Deprovisioning — the control that fails Type II most often.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The gap is the finding.&lt;/strong&gt; The dangerous window is between a termination date and the moment access is actually removed. Auditors sample terminated employees and check the revocation timestamp against the leave date.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automate it or lose it.&lt;/strong&gt; Manual deprovisioning drifts; an automated daily check that anti-joins live grants against the active roster catches the orphan before the auditor does.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjupqsfuyeq6vc5t5iizw.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjupqsfuyeq6vc5t5iizw.jpeg" alt="Iconographic SOC 2 access-control diagram — an HR roster driving joiner-mover-leaver provisioning, an RBAC role hierarchy granting least privilege, and an anti-join finding orphaned grants left after a termination date." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — a least-privilege role hierarchy in Snowflake
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The clearest demonstration of least privilege is a role hierarchy where table privileges live on functional roles and people inherit only what they need. No human holds a direct table grant; a person is granted a role, and the role holds the privilege. That indirection is what makes access reviewable and revocable in one statement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Grant an analyst read access to a &lt;code&gt;customers&lt;/code&gt; table but keep the &lt;code&gt;email&lt;/code&gt; column masked, using roles rather than direct user grants.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; A &lt;code&gt;customers&lt;/code&gt; table with an &lt;code&gt;email&lt;/code&gt; PII column, and a user &lt;code&gt;ada&lt;/code&gt; who should read it masked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- 1. Functional role holds the privilege, not the user.&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;ROLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;read_customers&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;GRANT&lt;/span&gt; &lt;span class="k"&gt;USAGE&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;DATABASE&lt;/span&gt; &lt;span class="n"&gt;analytics&lt;/span&gt; &lt;span class="k"&gt;TO&lt;/span&gt; &lt;span class="k"&gt;ROLE&lt;/span&gt; &lt;span class="n"&gt;read_customers&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;GRANT&lt;/span&gt; &lt;span class="k"&gt;USAGE&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;SCHEMA&lt;/span&gt; &lt;span class="n"&gt;analytics&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;core&lt;/span&gt; &lt;span class="k"&gt;TO&lt;/span&gt; &lt;span class="k"&gt;ROLE&lt;/span&gt; &lt;span class="n"&gt;read_customers&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;GRANT&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;analytics&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;core&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customers&lt;/span&gt; &lt;span class="k"&gt;TO&lt;/span&gt; &lt;span class="k"&gt;ROLE&lt;/span&gt; &lt;span class="n"&gt;read_customers&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- 2. Column masking policy — email is only unmasked for a privileged role.&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="n"&gt;MASKING&lt;/span&gt; &lt;span class="n"&gt;POLICY&lt;/span&gt; &lt;span class="n"&gt;mask_email&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;val&lt;/span&gt; &lt;span class="n"&gt;STRING&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;RETURNS&lt;/span&gt; &lt;span class="n"&gt;STRING&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
  &lt;span class="k"&gt;CASE&lt;/span&gt; &lt;span class="k"&gt;WHEN&lt;/span&gt; &lt;span class="k"&gt;CURRENT_ROLE&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'PII_ADMIN'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;THEN&lt;/span&gt; &lt;span class="n"&gt;val&lt;/span&gt;
       &lt;span class="k"&gt;ELSE&lt;/span&gt; &lt;span class="n"&gt;REGEXP_REPLACE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;val&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'.+@'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'****@'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;END&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;analytics&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;core&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customers&lt;/span&gt;
  &lt;span class="k"&gt;MODIFY&lt;/span&gt; &lt;span class="k"&gt;COLUMN&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;MASKING&lt;/span&gt; &lt;span class="n"&gt;POLICY&lt;/span&gt; &lt;span class="n"&gt;mask_email&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- 3. The person inherits the role — never a direct table grant.&lt;/span&gt;
&lt;span class="k"&gt;GRANT&lt;/span&gt; &lt;span class="k"&gt;ROLE&lt;/span&gt; &lt;span class="n"&gt;read_customers&lt;/span&gt; &lt;span class="k"&gt;TO&lt;/span&gt; &lt;span class="k"&gt;USER&lt;/span&gt; &lt;span class="k"&gt;ada&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Step 1 puts every privilege on &lt;code&gt;read_customers&lt;/code&gt;, so the grant is auditable as one row and revocable with one &lt;code&gt;REVOKE ROLE&lt;/code&gt;. Step 2 attaches a masking policy to the &lt;code&gt;email&lt;/code&gt; column: any role except &lt;code&gt;PII_ADMIN&lt;/code&gt; sees &lt;code&gt;****@domain&lt;/code&gt;, so the table grant does not leak the PII. Step 3 grants the &lt;em&gt;role&lt;/em&gt; to &lt;code&gt;ada&lt;/code&gt;; she never receives a direct table privilege, so a reviewer sees exactly one line — "ada has read_customers" — instead of hunting through per-object grants.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;principal&lt;/th&gt;
&lt;th&gt;can read customers&lt;/th&gt;
&lt;th&gt;sees raw email?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;ada&lt;/code&gt; (via &lt;code&gt;read_customers&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;no — masked &lt;code&gt;****@&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;PII_ADMIN&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;everyone else&lt;/td&gt;
&lt;td&gt;no grant&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If revoking a person's access takes more than one &lt;code&gt;REVOKE ROLE&lt;/code&gt;, your access model is not least-privilege — privileges belong on roles, and people belong in roles.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on deprovisioning
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Your auditor samples five terminated employees and asks you to prove none of them retained warehouse access after their leave date. Write a query that lists any active grant whose grantee is no longer an active employee — the orphaned-access anti-join.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using an anti-join against the active-employee roster
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- grants_snapshot: current warehouse grants (grantee_name, role, granted_on)&lt;/span&gt;
&lt;span class="c1"&gt;-- hr_roster: employees with status and termination_date&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="k"&gt;g&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;grantee_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;g&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;role&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;termination_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;DATEDIFF&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'day'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;termination_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;CURRENT_DATE&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;days_orphaned&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;grants_snapshot&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;g&lt;/span&gt;
&lt;span class="k"&gt;LEFT&lt;/span&gt; &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;hr_roster&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;
       &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;UPPER&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;g&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;grantee_name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;UPPER&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;warehouse_user&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;warehouse_user&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;          &lt;span class="c1"&gt;-- grant with no matching employee at all&lt;/span&gt;
   &lt;span class="k"&gt;OR&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'terminated'&lt;/span&gt;           &lt;span class="c1"&gt;-- or a terminated employee still granted&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;days_orphaned&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;grantee&lt;/th&gt;
&lt;th&gt;in roster?&lt;/th&gt;
&lt;th&gt;status&lt;/th&gt;
&lt;th&gt;verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ADA&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;active&lt;/td&gt;
&lt;td&gt;not returned (compliant)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LINUS&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;terminated 2026-02-10&lt;/td&gt;
&lt;td&gt;returned — orphaned 33 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SVC_ETL&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;active (service)&lt;/td&gt;
&lt;td&gt;not returned&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GHOST&lt;/td&gt;
&lt;td&gt;no match&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;returned — grant with no owner&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Every current grant is the driving side of a &lt;code&gt;LEFT JOIN&lt;/code&gt; onto the HR roster keyed on the warehouse username.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;WHERE r.warehouse_user IS NULL&lt;/code&gt; catches grants whose grantee is not in the roster at all — the un-owned accounts auditors hate.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;r.status = 'terminated'&lt;/code&gt; catches grantees who are in the roster but have left — the deprovisioning gap.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;days_orphaned&lt;/code&gt; quantifies exposure so the worst offenders sort to the top and the remediation is prioritized.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;grantee_name&lt;/th&gt;
&lt;th&gt;role&lt;/th&gt;
&lt;th&gt;termination_date&lt;/th&gt;
&lt;th&gt;days_orphaned&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;LINUS&lt;/td&gt;
&lt;td&gt;read_customers&lt;/td&gt;
&lt;td&gt;2026-02-10&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GHOST&lt;/td&gt;
&lt;td&gt;loader&lt;/td&gt;
&lt;td&gt;(none)&lt;/td&gt;
&lt;td&gt;(unknown owner)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Anti-join&lt;/strong&gt;&lt;/strong&gt; — a &lt;code&gt;LEFT JOIN&lt;/code&gt; plus &lt;code&gt;WHERE right IS NULL&lt;/code&gt; (widened here to include terminated rows) returns exactly the grants that &lt;em&gt;should not&lt;/em&gt; exist; it is the canonical "find what is missing from the allowed set" pattern.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Authoritative roster&lt;/strong&gt;&lt;/strong&gt; — the HR/IdP roster is the source of truth for "who is an employee," so access correctness is defined against it, not against a hand-maintained list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Deprovisioning gap&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;days_orphaned&lt;/code&gt; turns a boolean finding into a measured risk window, which is exactly the metric a Type II auditor tests against the leave date.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Service accounts&lt;/strong&gt;&lt;/strong&gt; — un-matched grantees surface non-human accounts that need a named owner, closing the "who owns GHOST?" question before it becomes a finding.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — O(grants) with a hash join on the roster; a nightly run over a few thousand grants is milliseconds and produces continuous deprovisioning evidence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Access&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — access-control&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Least-privilege and access-review problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Quality&lt;/span&gt;
&lt;span&gt;Topic — data-validation&lt;/span&gt;
&lt;strong&gt;Roster-reconciliation and validation problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-validation" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. Warehouse audit logs &amp;amp; access history
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Query and access history are the evidence that access controls actually held — logging is what turns a policy into proof
&lt;/h3&gt;

&lt;p&gt;An access policy you cannot observe is a policy you cannot attest. The Security criterion's monitoring controls demand that you can answer "who read this, who changed that, when, and from where" for the whole audit period. Modern warehouses hand you this for free through account-usage views — the trick is knowing which view proves which control, and respecting their latency and retention. Say it plainly: &lt;strong&gt;audit logs are the operating-effectiveness evidence; without them a well-designed control still fails a Type II&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The log surfaces you actually cite.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;LOGIN_HISTORY&lt;/code&gt;.&lt;/strong&gt; Every authentication attempt, success or failure, with client IP and method. Proves the authentication and MFA-enforcement controls and surfaces brute-force patterns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;QUERY_HISTORY&lt;/code&gt;.&lt;/strong&gt; Every statement run, by whom, against what, with duration and bytes scanned. The backbone of "who ran what" evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ACCESS_HISTORY&lt;/code&gt;.&lt;/strong&gt; The strong one for confidentiality: per-query, the exact objects and &lt;em&gt;columns&lt;/em&gt; read and written. This is what proves who touched a PII column, not just who queried the table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;GRANTS_TO_USERS&lt;/code&gt; / &lt;code&gt;GRANTS_TO_ROLES&lt;/code&gt;.&lt;/strong&gt; Point-in-time snapshots of the access model itself — the evidence behind the access-review control from section 2.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Column-level access is the confidentiality proof.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Table-level is not enough.&lt;/strong&gt; "User X selected from &lt;code&gt;customers&lt;/code&gt;" does not prove whether they read the masked &lt;code&gt;email&lt;/code&gt;. &lt;code&gt;ACCESS_HISTORY&lt;/code&gt; records the columns actually referenced, so you can prove column-level confidentiality.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Masking and row-access policies leave a trail too.&lt;/strong&gt; Applying and altering a policy is itself a change auditors can see, tying the confidentiality control back to change management.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Latency and retention — the trap.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Account-usage views lag.&lt;/strong&gt; Snowflake's &lt;code&gt;ACCOUNT_USAGE&lt;/code&gt; views can be delayed (often up to ~2–3 hours for access history), so a real-time control cannot depend on them; an evidence-collection job must account for the lag.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retention is finite.&lt;/strong&gt; &lt;code&gt;ACCOUNT_USAGE&lt;/code&gt; retains history for a bounded period (on the order of a year for many views, less for some). A Type II window can exceed retention, so you must &lt;em&gt;export&lt;/em&gt; logs to durable storage or you will have gaps exactly when the auditor samples an old date.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Other platforms, same idea.&lt;/strong&gt; BigQuery exposes Cloud Audit Logs (Admin Activity and Data Access) and &lt;code&gt;INFORMATION_SCHEMA.JOBS&lt;/code&gt;; Redshift has &lt;code&gt;STL&lt;/code&gt;/&lt;code&gt;SVL&lt;/code&gt; system tables and CloudTrail. The pattern is identical: system views prove access, and you export them before they age out.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3inllhxd4tr29n2140uh.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3inllhxd4tr29n2140uh.jpeg" alt="Iconographic warehouse audit-log diagram — QUERY_HISTORY, LOGIN_HISTORY and ACCESS_HISTORY views feeding a column-level access record that shows which principal read a PII column, with a masking-policy glyph." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — reading who queried a sensitive table
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday audit query answers "show me every access to this table in the review window." It is the first thing you run when an auditor points at a sensitive object, and the shape — filter &lt;code&gt;QUERY_HISTORY&lt;/code&gt; by object and time — is the foundation the harder column-level query builds on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; List every user who queried &lt;code&gt;analytics.core.customers&lt;/code&gt; in the last 90 days, with how many times and when they last touched it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; &lt;code&gt;QUERY_HISTORY&lt;/code&gt; rows over the last 90 days referencing various tables.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;user_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;              &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;query_count&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;MAX&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start_time&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;       &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;last_accessed&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;snowflake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;query_history&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;start_time&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;DATEADD&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'day'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;90&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;CURRENT_TIMESTAMP&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;query_text&lt;/span&gt; &lt;span class="k"&gt;ILIKE&lt;/span&gt; &lt;span class="s1"&gt;'%analytics.core.customers%'&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;user_name&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;query_count&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The &lt;code&gt;WHERE&lt;/code&gt; clause bounds the window to the 90-day review period and filters to statements referencing the target table. &lt;code&gt;GROUP BY user_name&lt;/code&gt; collapses each principal to one row; &lt;code&gt;COUNT(*)&lt;/code&gt; is their access frequency and &lt;code&gt;MAX(start_time)&lt;/code&gt; is the recency an auditor asks for. Filtering on &lt;code&gt;query_text&lt;/code&gt; is the quick approximation — the precise, column-aware version uses &lt;code&gt;ACCESS_HISTORY&lt;/code&gt;, which the interview question below demands.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;user_name&lt;/th&gt;
&lt;th&gt;query_count&lt;/th&gt;
&lt;th&gt;last_accessed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ADA&lt;/td&gt;
&lt;td&gt;42&lt;/td&gt;
&lt;td&gt;2026-09-12 08:15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BI_SERVICE&lt;/td&gt;
&lt;td&gt;310&lt;/td&gt;
&lt;td&gt;2026-09-14 23:59&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LINUS&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;2026-02-09 17:40&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; &lt;code&gt;QUERY_HISTORY&lt;/code&gt; proves &lt;em&gt;that a table was queried&lt;/em&gt;; when the question is &lt;em&gt;which columns were read&lt;/em&gt;, you must move to &lt;code&gt;ACCESS_HISTORY&lt;/code&gt; — text matching cannot prove column-level confidentiality.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on column-level access history
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; An auditor wants proof of exactly which principals read the &lt;code&gt;email&lt;/code&gt; PII column of &lt;code&gt;customers&lt;/code&gt; in the last 90 days. Table-level history is not enough. Write the query using Snowflake &lt;code&gt;ACCESS_HISTORY&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using LATERAL FLATTEN over ACCESS_HISTORY
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;ah&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;user_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;cols&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;columnName&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;STRING&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;column_read&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                      &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;reads&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;MAX&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ah&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;query_start_time&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;last_read&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;snowflake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;access_history&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;ah&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
     &lt;span class="k"&gt;LATERAL&lt;/span&gt; &lt;span class="n"&gt;FLATTEN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;input&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;ah&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;base_objects_accessed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
     &lt;span class="k"&gt;LATERAL&lt;/span&gt; &lt;span class="n"&gt;FLATTEN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;input&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;        &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;cols&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;ah&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;query_start_time&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;DATEADD&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'day'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;90&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;CURRENT_TIMESTAMP&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;objectName&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;STRING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'ANALYTICS.CORE.CUSTOMERS'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;cols&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;columnName&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;STRING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'EMAIL'&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;ah&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;user_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;column_read&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="k"&gt;reads&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;access_history row&lt;/th&gt;
&lt;th&gt;object&lt;/th&gt;
&lt;th&gt;columns array&lt;/th&gt;
&lt;th&gt;matches email?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;q1 by ADA&lt;/td&gt;
&lt;td&gt;CUSTOMERS&lt;/td&gt;
&lt;td&gt;[id, name, email]&lt;/td&gt;
&lt;td&gt;yes — 1 read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;q2 by ADA&lt;/td&gt;
&lt;td&gt;CUSTOMERS&lt;/td&gt;
&lt;td&gt;[id, name]&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;q3 by BI_SERVICE&lt;/td&gt;
&lt;td&gt;CUSTOMERS&lt;/td&gt;
&lt;td&gt;[id, email]&lt;/td&gt;
&lt;td&gt;yes — 1 read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;q4 by LINUS&lt;/td&gt;
&lt;td&gt;ORDERS&lt;/td&gt;
&lt;td&gt;[order_id]&lt;/td&gt;
&lt;td&gt;no — wrong object&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;ACCESS_HISTORY.base_objects_accessed&lt;/code&gt; is a semi-structured array of the base objects each query touched; the first &lt;code&gt;FLATTEN&lt;/code&gt; yields one row per object.&lt;/li&gt;
&lt;li&gt;Each object carries a nested &lt;code&gt;columns&lt;/code&gt; array; the second &lt;code&gt;FLATTEN&lt;/code&gt; yields one row per column actually referenced — this is the column-level granularity table-level history cannot give.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;WHERE&lt;/code&gt; filters to the target object and the &lt;code&gt;EMAIL&lt;/code&gt; column, so only genuine reads of the PII column survive.&lt;/li&gt;
&lt;li&gt;Grouping by user and column produces per-principal read counts and recency — precisely the confidentiality evidence the auditor asked for.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;user_name&lt;/th&gt;
&lt;th&gt;column_read&lt;/th&gt;
&lt;th&gt;reads&lt;/th&gt;
&lt;th&gt;last_read&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;BI_SERVICE&lt;/td&gt;
&lt;td&gt;EMAIL&lt;/td&gt;
&lt;td&gt;118&lt;/td&gt;
&lt;td&gt;2026-09-14 23:59&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ADA&lt;/td&gt;
&lt;td&gt;EMAIL&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;2026-09-11 10:02&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;base_objects_accessed&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;ACCESS_HISTORY&lt;/code&gt; records the &lt;em&gt;resolved&lt;/em&gt; base objects and columns per query, so it sees through views and &lt;code&gt;SELECT *&lt;/code&gt; to the real columns read — the only reliable source for column-level proof.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;LATERAL FLATTEN&lt;/strong&gt;&lt;/strong&gt; — flattening the nested object and column arrays turns one query row into one row per column, letting a normal &lt;code&gt;GROUP BY&lt;/code&gt; count column-level access.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Column-level confidentiality&lt;/strong&gt;&lt;/strong&gt; — the result names the exact principals who read &lt;code&gt;email&lt;/code&gt;, which is what the Confidentiality criterion requires and what table-level history can only approximate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Export before retention&lt;/strong&gt;&lt;/strong&gt; — because &lt;code&gt;ACCESS_HISTORY&lt;/code&gt; ages out, this query must run inside an evidence job that persists results, or the proof vanishes before the audit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — O(rows × columns-per-query) over the flattened set; scanning a bounded 90-day window keeps it cheap, and pushing the object filter down limits the scan.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Access&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — access-control&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Column-level access and audit-trail problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Audit&lt;/span&gt;
&lt;span&gt;Topic — event-log&lt;/span&gt;
&lt;strong&gt;Event-log and access-history query problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/event-log" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Change management for pipelines
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Every pipeline change must trace to a reviewed, approved, controlled deploy — CC8 is the control auditors reconstruct from your git and CI history
&lt;/h3&gt;

&lt;p&gt;The Security criterion's change-management control (&lt;code&gt;CC8.1&lt;/code&gt;) asks a simple question with expensive consequences: can you prove that every change to production went through review and approval? For a data team, "production" is the pipeline code, the transformations, and the schema. The auditor reconstructs the control from version control and CI, so the control is only as good as the trail those systems leave. Say it in one line: &lt;strong&gt;an approved, peer-reviewed pull request is the unit of controlled change, and a prod deploy without one is the finding&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The change-control chain.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Version control is the system of record.&lt;/strong&gt; All pipeline and transformation code lives in git; nothing reaches production except through a merge. A change with no commit is an un-auditable change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Peer review is the control.&lt;/strong&gt; A pull request reviewed and approved by someone other than the author is the evidence. The approval, the reviewer identity, and the timestamp are the artifact.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CI/CD gates the deploy.&lt;/strong&gt; The deployment pipeline should refuse to ship a branch that was not merged through an approved PR, and it should record what it shipped — commit SHA, PR number, actor, time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Segregation of duties.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Author is not approver.&lt;/strong&gt; The person who wrote the change cannot be the one who approves it — the review is meaningless otherwise, and auditors check for self-approval.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Approver is not necessarily deployer, and deploys are automated.&lt;/strong&gt; Manual, un-logged production access is the anti-pattern; an automated deploy from an approved merge removes the human who could bypass the control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Break-glass is logged.&lt;/strong&gt; Emergency changes happen; the control is not "never," it is "every emergency change is logged, justified, and reviewed after the fact."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The evidence is a join.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A deployments table.&lt;/strong&gt; Each production deploy records &lt;code&gt;deploy_id&lt;/code&gt;, &lt;code&gt;commit_sha&lt;/code&gt;, &lt;code&gt;pr_number&lt;/code&gt;, &lt;code&gt;deployed_by&lt;/code&gt;, &lt;code&gt;deployed_at&lt;/code&gt;. This is your change log.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An approvals table.&lt;/strong&gt; Each merged PR records &lt;code&gt;pr_number&lt;/code&gt;, &lt;code&gt;author&lt;/code&gt;, &lt;code&gt;approved_by&lt;/code&gt;, &lt;code&gt;approved_at&lt;/code&gt;. Pulled from the git host's API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The control test is the reconciliation.&lt;/strong&gt; Left-join deploys to approvals; any deploy with no approved PR, or where approver equals author, is a change-control exception — the exact thing a Type II samples.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5szbv3rsd18al2lvz4bk.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5szbv3rsd18al2lvz4bk.jpeg" alt="Iconographic change-management diagram — a peer-reviewed pull request with an approval gate feeding a CI deploy, a deployments table joined to approvals, and a segregation-of-duties note that author, approver, and deployer differ." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — recording a deploy with its approval lineage
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The control produces evidence only if the deploy step &lt;em&gt;writes it down&lt;/em&gt;. The cleanest pattern is a deploy job that, at ship time, records the commit, the PR it came from, and who triggered it into a deployments table. That single insert is what a whole quarter of change-control evidence is built on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; At deploy time, capture the change-control lineage of a pipeline release into a &lt;code&gt;deployments&lt;/code&gt; table so it can later be reconciled against approvals.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; A release of commit &lt;code&gt;9f3a1c&lt;/code&gt; merged via PR &lt;code&gt;412&lt;/code&gt;, deployed by the CI service account.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;dt&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;record_deploy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;commit_sha&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pr_number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;actor&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;row&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deploy_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sha1&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;commit_sha&lt;/span&gt;&lt;span class="si"&gt;}{&lt;/span&gt;&lt;span class="n"&gt;pr_number&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()[:&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;commit_sha&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;commit_sha&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pr_number&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;pr_number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deployed_by&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;actor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deployed_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;dt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;utcnow&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;isoformat&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;environment&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prod&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;INSERT INTO change_control.deployments
           (deploy_id, commit_sha, pr_number, deployed_by, deployed_at, environment)
           VALUES (%(deploy_id)s, %(commit_sha)s, %(pr_number)s,
                   %(deployed_by)s, %(deployed_at)s, %(environment)s)&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt;

&lt;span class="nf"&gt;record_deploy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;commit_sha&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;9f3a1c&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pr_number&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;412&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;actor&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ci-deployer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# CI-only, never a human shell
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The deploy job — not a human — calls &lt;code&gt;record_deploy&lt;/code&gt; as its final step, so the evidence is a byproduct of shipping, impossible to forget. It captures the &lt;code&gt;commit_sha&lt;/code&gt; and the &lt;code&gt;pr_number&lt;/code&gt; that carried it, plus the acting identity and a UTC timestamp. Because CI is the only path to prod, the deployments table becomes a complete census of production changes, and the &lt;code&gt;pr_number&lt;/code&gt; is the foreign key that lets the auditor's reconciliation join to the approvals table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;deploy_id&lt;/th&gt;
&lt;th&gt;commit_sha&lt;/th&gt;
&lt;th&gt;pr_number&lt;/th&gt;
&lt;th&gt;deployed_by&lt;/th&gt;
&lt;th&gt;environment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;3d9e1a77c0b2&lt;/td&gt;
&lt;td&gt;9f3a1c&lt;/td&gt;
&lt;td&gt;412&lt;/td&gt;
&lt;td&gt;ci-deployer&lt;/td&gt;
&lt;td&gt;prod&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If a deploy can happen without writing a row to the deployments table, your change-control evidence has holes — make the evidence write a mandatory step of the deploy, not a manual afterthought.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on change-control exceptions
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Prove that every production deploy in the audit window came from a peer-approved PR, and that no change was self-approved. Write the query that returns the change-control exceptions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using an anti-join with a segregation-of-duties check
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deploy_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pr_number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deployed_by&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;author&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;approved_by&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;CASE&lt;/span&gt;
        &lt;span class="k"&gt;WHEN&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pr_number&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;              &lt;span class="k"&gt;THEN&lt;/span&gt; &lt;span class="s1"&gt;'no approved PR'&lt;/span&gt;
        &lt;span class="k"&gt;WHEN&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;approved_by&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;author&lt;/span&gt;         &lt;span class="k"&gt;THEN&lt;/span&gt; &lt;span class="s1"&gt;'self-approved'&lt;/span&gt;
    &lt;span class="k"&gt;END&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;exception_type&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;change_control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deployments&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;
&lt;span class="k"&gt;LEFT&lt;/span&gt; &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;change_control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;approved_prs&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;
       &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pr_number&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pr_number&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environment&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'prod'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deployed_at&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-01'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pr_number&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;OR&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;approved_by&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;author&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deployed_at&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;deploy&lt;/th&gt;
&lt;th&gt;pr_number&lt;/th&gt;
&lt;th&gt;approved?&lt;/th&gt;
&lt;th&gt;author=approver?&lt;/th&gt;
&lt;th&gt;verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;dep_1&lt;/td&gt;
&lt;td&gt;412&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;compliant — not returned&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;dep_2&lt;/td&gt;
&lt;td&gt;419&lt;/td&gt;
&lt;td&gt;none in approvals&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;returned — no approved PR&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;dep_3&lt;/td&gt;
&lt;td&gt;421&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;ada = ada&lt;/td&gt;
&lt;td&gt;returned — self-approved&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;dep_4&lt;/td&gt;
&lt;td&gt;430&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;compliant — not returned&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Every prod deploy in the window is the driving side of a &lt;code&gt;LEFT JOIN&lt;/code&gt; onto approved PRs, keyed on &lt;code&gt;pr_number&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;a.pr_number IS NULL&lt;/code&gt; catches deploys whose PR was never approved (or never existed) — an uncontrolled change.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;a.approved_by = a.author&lt;/code&gt; catches the segregation-of-duties violation where the author approved their own change.&lt;/li&gt;
&lt;li&gt;Only exceptions survive the &lt;code&gt;WHERE&lt;/code&gt;; an empty result set is the passing evidence, and any rows are the exact deploys the auditor will drill into.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;deploy_id&lt;/th&gt;
&lt;th&gt;pr_number&lt;/th&gt;
&lt;th&gt;exception_type&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;dep_2&lt;/td&gt;
&lt;td&gt;419&lt;/td&gt;
&lt;td&gt;no approved PR&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;dep_3&lt;/td&gt;
&lt;td&gt;421&lt;/td&gt;
&lt;td&gt;self-approved&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Change-control chain&lt;/strong&gt;&lt;/strong&gt; — modeling deploys and approvals as two joinable tables turns a fuzzy policy ("we review changes") into a testable reconciliation with a boolean outcome per deploy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Anti-join for missing approvals&lt;/strong&gt;&lt;/strong&gt; — the &lt;code&gt;LEFT JOIN ... IS NULL&lt;/code&gt; pattern isolates deploys with no approval, the same shape as the deprovisioning check, reused for a different control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Segregation of duties&lt;/strong&gt;&lt;/strong&gt; — comparing &lt;code&gt;approved_by&lt;/code&gt; to &lt;code&gt;author&lt;/code&gt; encodes the "author is not approver" rule directly in SQL, so self-approval cannot hide.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Empty-set-is-passing&lt;/strong&gt;&lt;/strong&gt; — a control whose passing state is "zero exceptions" is easy to monitor continuously and easy to alert on the moment a violation appears.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — O(deploys) with a hash join on &lt;code&gt;pr_number&lt;/code&gt;; trivial to run per deploy as a gate and nightly as evidence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Reliability&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — reliability&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Change-control and deployment-safety problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/reliability" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;SLA&lt;/span&gt;
&lt;span&gt;Topic — sla-monitoring&lt;/span&gt;
&lt;strong&gt;Monitoring and control-status problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/sla-monitoring" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Evidence-collection pipelines
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Automated evidence is what makes a Type II survivable — a pipeline that snapshots controls into an immutable, tamper-evident ledger
&lt;/h3&gt;

&lt;p&gt;A Type II audit does not sample your controls on one lucky day; it samples them across months, and it expects the evidence to have existed &lt;em&gt;at the time&lt;/em&gt;, not to be reconstructed afterward. Collecting that evidence by hand — screenshots, spreadsheets, quarterly scrambles — is where audits go to die. The senior move is to treat evidence like any other pipeline output: &lt;strong&gt;scheduled jobs snapshot each control's evidence into an append-only, tamper-evident store, so the audit is a query, not a project&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What an evidence-collection pipeline gathers.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Access snapshots.&lt;/strong&gt; Nightly captures of &lt;code&gt;GRANTS_TO_USERS&lt;/code&gt; / &lt;code&gt;GRANTS_TO_ROLES&lt;/code&gt; so every point-in-time access state exists for the whole window, not just today.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Access-review artifacts.&lt;/strong&gt; The quarterly review's input snapshot, the reviewer, and the resulting revocations, stored together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log exports.&lt;/strong&gt; &lt;code&gt;ACCESS_HISTORY&lt;/code&gt;, &lt;code&gt;QUERY_HISTORY&lt;/code&gt;, and &lt;code&gt;LOGIN_HISTORY&lt;/code&gt; exported before they age out of retention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Control-test results.&lt;/strong&gt; The anti-join outputs from sections 2–4 (orphaned grants, PII reads, change-control exceptions), each run and stored with its timestamp so "the control ran and passed on this date" is itself evidence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Immutability — what makes evidence trustworthy.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Append-only, never update.&lt;/strong&gt; Evidence tables are insert-only. An evidence store you can &lt;code&gt;UPDATE&lt;/code&gt; or &lt;code&gt;DELETE&lt;/code&gt; is one an auditor cannot trust; the control is only credible if the record cannot be quietly changed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hash-chaining for tamper-evidence.&lt;/strong&gt; Each row stores a hash of its own contents plus the previous row's hash. Altering any historical row breaks the chain from that point forward, so tampering is detectable without trusting the database alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WORM / Object Lock storage.&lt;/strong&gt; Exported evidence lands in write-once-read-many object storage (S3 Object Lock in compliance mode, GCS retention lock) so even an administrator cannot delete it before its retention expires.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Continuous control monitoring.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A control that reports its own status.&lt;/strong&gt; Instead of proving controls at audit time, each control emits a &lt;code&gt;PASS&lt;/code&gt;/&lt;code&gt;FAIL&lt;/code&gt; with its evidence on every run, feeding a monitoring dashboard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alert on transition, not on state.&lt;/strong&gt; The signal that matters is a control flipping from &lt;code&gt;PASS&lt;/code&gt; to &lt;code&gt;FAIL&lt;/code&gt; — an orphaned grant appearing, a self-approved deploy landing — which routes to on-call the same day, shrinking the exposure window the auditor would otherwise measure.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwsj3495ghmeut12zvu1d.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwsj3495ghmeut12zvu1d.jpeg" alt="Iconographic evidence-pipeline diagram — a scheduled job snapshotting grants, access reviews, and logs into an append-only hash-chained audit table stored on WORM object-lock storage, with a control-monitoring status glyph." width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — an append-only evidence snapshot
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The building block of an evidence pipeline is a job that runs a control's evidence query and &lt;em&gt;appends&lt;/em&gt; the result with a run timestamp — never overwriting the previous run. Keeping every run is what gives you a point-in-time record for any date the auditor picks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Snapshot today's grant listing into an append-only &lt;code&gt;evidence_grants&lt;/code&gt; table so that every past day's access state is preserved for the audit window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; The current output of the grants query, captured on run date &lt;code&gt;2026-09-15&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;dt&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;snapshot_grants&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;run_ts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;utcnow&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;isoformat&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;INSERT INTO evidence.evidence_grants
           (captured_at, grantee_name, role, privilege)
           SELECT %(ts)s, grantee_name, role, privilege
           FROM   snowflake.account_usage.grants_to_users
           WHERE  deleted_on IS NULL&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;     &lt;span class="c1"&gt;# only currently-live grants
&lt;/span&gt;        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;run_ts&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# NOTE: no UPDATE / DELETE anywhere — the table is insert-only.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The job stamps every captured row with a single &lt;code&gt;captured_at&lt;/code&gt; for the run, then inserts the live grant set as of that moment. Because the table is only ever inserted into, run &lt;em&gt;N+1&lt;/em&gt; never disturbs run &lt;em&gt;N&lt;/em&gt; — the history of "who had what access on any given day" accumulates. Filtering &lt;code&gt;deleted_on IS NULL&lt;/code&gt; captures the live state; the accumulation of daily snapshots is what lets the auditor ask "show me access on 2026-06-30" months later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;captured_at&lt;/th&gt;
&lt;th&gt;grantee_name&lt;/th&gt;
&lt;th&gt;role&lt;/th&gt;
&lt;th&gt;privilege&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-15T00:00&lt;/td&gt;
&lt;td&gt;ADA&lt;/td&gt;
&lt;td&gt;read_customers&lt;/td&gt;
&lt;td&gt;SELECT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-15T00:00&lt;/td&gt;
&lt;td&gt;BI_SERVICE&lt;/td&gt;
&lt;td&gt;read_raw&lt;/td&gt;
&lt;td&gt;SELECT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-14T00:00&lt;/td&gt;
&lt;td&gt;ADA&lt;/td&gt;
&lt;td&gt;read_customers&lt;/td&gt;
&lt;td&gt;SELECT&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Evidence tables grow forever and change never — the day you find yourself wanting to &lt;code&gt;UPDATE&lt;/code&gt; an evidence row is the day your evidence stops being evidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Python interview question on tamper-evident evidence
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; An auditor asks how you would prove your evidence ledger has not been altered after the fact. Design an append-only evidence table where any tampering with a historical row is detectable, and show how the check works.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using a hash-chained immutable ledger
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;row_hash&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prev_hash&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;prev_hash&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sort_keys&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;append_evidence&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ledger&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;prev&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ledger&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;row_hash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ledger&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GENESIS&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;entry&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;seq&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ledger&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;payload&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prev_hash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;row_hash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;row_hash&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;
    &lt;span class="n"&gt;ledger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                 &lt;span class="c1"&gt;# append-only
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ledger&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;prev&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GENESIS&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;ledger&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prev_hash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;prev&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;row_hash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="nf"&gt;row_hash&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;payload&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;seq&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;              &lt;span class="c1"&gt;# first tampered sequence, or None if intact
&lt;/span&gt;        &lt;span class="n"&gt;prev&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;row_hash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;seq&lt;/th&gt;
&lt;th&gt;payload&lt;/th&gt;
&lt;th&gt;prev_hash&lt;/th&gt;
&lt;th&gt;row_hash&lt;/th&gt;
&lt;th&gt;chain ok?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;grants@day1&lt;/td&gt;
&lt;td&gt;GENESIS&lt;/td&gt;
&lt;td&gt;h0&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;grants@day2&lt;/td&gt;
&lt;td&gt;h0&lt;/td&gt;
&lt;td&gt;h1&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;grants@day3 (tampered)&lt;/td&gt;
&lt;td&gt;h1&lt;/td&gt;
&lt;td&gt;h2' ≠ recompute&lt;/td&gt;
&lt;td&gt;breaks here&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;grants@day4&lt;/td&gt;
&lt;td&gt;h2 (stale)&lt;/td&gt;
&lt;td&gt;h3&lt;/td&gt;
&lt;td&gt;fails — prev mismatch&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Each entry's &lt;code&gt;row_hash&lt;/code&gt; is &lt;code&gt;sha256(prev_hash + payload)&lt;/code&gt;, so every row cryptographically depends on the entire history before it.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;append_evidence&lt;/code&gt; only ever appends and always chains onto the last row's hash — there is no code path that rewrites a prior entry.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;verify&lt;/code&gt; recomputes each hash from the previous one; if an attacker edits &lt;code&gt;day3&lt;/code&gt;'s payload, its recomputed hash no longer matches the stored &lt;code&gt;row_hash&lt;/code&gt;, and every subsequent link fails too.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;verify&lt;/code&gt; returns the first broken &lt;code&gt;seq&lt;/code&gt;, pinpointing exactly where the ledger was tampered — evidence that the evidence itself is intact.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;ledger state&lt;/th&gt;
&lt;th&gt;verify() returns&lt;/th&gt;
&lt;th&gt;meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;untouched&lt;/td&gt;
&lt;td&gt;&lt;code&gt;None&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;chain intact — evidence trustworthy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;row 2 edited&lt;/td&gt;
&lt;td&gt;&lt;code&gt;2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;tamper detected at seq 2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Append-only ledger&lt;/strong&gt;&lt;/strong&gt; — insert-only semantics remove the "quiet update" attack entirely; combined with WORM storage, even a privileged operator cannot rewrite history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Hash chaining&lt;/strong&gt;&lt;/strong&gt; — binding each row to the previous hash makes the ledger tamper-evident: changing any row invalidates every row after it, so a single stored final hash attests the whole chain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Deterministic serialization&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;json.dumps(..., sort_keys=True)&lt;/code&gt; guarantees the same payload always hashes identically, so a legitimate re-verification never produces a false tamper alarm.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Continuous verification&lt;/strong&gt;&lt;/strong&gt; — running &lt;code&gt;verify&lt;/code&gt; on a schedule turns tamper-evidence into tamper-&lt;em&gt;detection&lt;/em&gt;, alerting the day the chain breaks rather than at audit time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — append is O(1) and verification is O(n) over the ledger; a nightly full verify of a year of daily snapshots is a few hundred hashes, negligible against the audit assurance it buys.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;SLA&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — sla-monitoring&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Control-monitoring and alerting problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/sla-monitoring" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Quality&lt;/span&gt;
&lt;span&gt;Topic — data-quality&lt;/span&gt;
&lt;strong&gt;Immutable-ledger and integrity-check problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-quality" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — SOC 2 evidence recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;List all warehouse grants (access evidence).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;grantee_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;role&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;privilege&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;granted_on&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;snowflake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;grants_to_users&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="n"&gt;deleted_on&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;grantee_name&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Quarterly access review (diff against last snapshot).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;grantee_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;role&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;evidence_grants&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="n"&gt;captured_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;MAX&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;captured_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;evidence_grants&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;EXCEPT&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;grantee_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;role&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;evidence_grants&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="n"&gt;captured_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;previous_review_snapshot&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;-- rows added since last review&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Column-level PII access (confidentiality evidence).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;ah&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;user_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;columnName&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;STRING&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;reads&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;snowflake&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;access_history&lt;/span&gt; &lt;span class="n"&gt;ah&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;LATERAL&lt;/span&gt; &lt;span class="n"&gt;FLATTEN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;input&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;ah&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;base_objects_accessed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;LATERAL&lt;/span&gt; &lt;span class="n"&gt;FLATTEN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;input&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;objectName&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;STRING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'ANALYTICS.CORE.CUSTOMERS'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt;  &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;columnName&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;STRING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'EMAIL'&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Change-control exceptions (CC8 evidence).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deploy_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pr_number&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;change_control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deployments&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;
&lt;span class="k"&gt;LEFT&lt;/span&gt; &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;change_control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;approved_prs&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pr_number&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pr_number&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environment&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'prod'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pr_number&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;OR&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;approved_by&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;author&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;   &lt;span class="c1"&gt;-- empty = passing&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Immutable evidence table (insert-only, Object Lock target).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;evidence_grants&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;captured_at&lt;/span&gt;   &lt;span class="n"&gt;TIMESTAMP_NTZ&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;grantee_name&lt;/span&gt;  &lt;span class="n"&gt;STRING&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;role&lt;/span&gt;          &lt;span class="n"&gt;STRING&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;privilege&lt;/span&gt;     &lt;span class="n"&gt;STRING&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="c1"&gt;-- Grant only INSERT + SELECT to the evidence role; never UPDATE/DELETE.&lt;/span&gt;
&lt;span class="k"&gt;GRANT&lt;/span&gt; &lt;span class="k"&gt;INSERT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;evidence_grants&lt;/span&gt; &lt;span class="k"&gt;TO&lt;/span&gt; &lt;span class="k"&gt;ROLE&lt;/span&gt; &lt;span class="n"&gt;evidence_writer&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Trust criterion -&amp;gt; control -&amp;gt; evidence.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Trust criterion&lt;/th&gt;
&lt;th&gt;Control a DE owns&lt;/th&gt;
&lt;th&gt;Evidence artifact&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Security (CC6)&lt;/td&gt;
&lt;td&gt;Least privilege + deprovisioning&lt;/td&gt;
&lt;td&gt;grants snapshot + orphan anti-join&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security (CC7)&lt;/td&gt;
&lt;td&gt;Monitoring / audit logging&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ACCESS_HISTORY&lt;/code&gt; column-read export&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security (CC8)&lt;/td&gt;
&lt;td&gt;Change management&lt;/td&gt;
&lt;td&gt;deploy ⋈ approved-PR reconciliation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Confidentiality&lt;/td&gt;
&lt;td&gt;PII access restriction&lt;/td&gt;
&lt;td&gt;masked columns + column-access log&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Processing Integrity&lt;/td&gt;
&lt;td&gt;Pipeline correctness&lt;/td&gt;
&lt;td&gt;data-quality check results, stored&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is SOC 2 and which parts do data engineers own?
&lt;/h3&gt;

&lt;p&gt;SOC 2 is an independent attestation, performed by a licensed CPA firm against the AICPA Trust Services Criteria, that an organization's controls are suitably designed and — in a Type II — operated effectively over a period. It is a report you share with customers, not a certification you pass once. Data engineers own the controls that live in the data platform: logical access (least privilege, provisioning, deprovisioning), audit logging, change management on pipeline code, encryption of data at rest and in transit, and the evidence-collection that proves all of them.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between SOC 2 Type I and Type II?
&lt;/h3&gt;

&lt;p&gt;A Type I report opines that controls are suitably &lt;em&gt;designed&lt;/em&gt; as of a single point in time — a snapshot that answers "do the right controls exist?" A Type II report opines that those controls &lt;em&gt;operated effectively&lt;/em&gt; across a window, typically 3 to 12 months, and is tested by sampling evidence throughout that period. Type II is the one customers usually require, and it is the reason you must retain evidence continuously rather than assembling it at audit time.&lt;/p&gt;

&lt;h3&gt;
  
  
  What are the five Trust Services Criteria?
&lt;/h3&gt;

&lt;p&gt;They are Security, Availability, Processing Integrity, Confidentiality, and Privacy. Security — the Common Criteria — is mandatory in every SOC 2 and covers logical access, change management, risk assessment, and monitoring. The other four are included only if they are in scope for what you promise customers: Availability for uptime and recovery, Processing Integrity for complete and accurate processing, Confidentiality for protecting designated confidential data, and Privacy for handling personal information per your notice.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which warehouse audit logs prove access controls?
&lt;/h3&gt;

&lt;p&gt;In Snowflake, &lt;code&gt;LOGIN_HISTORY&lt;/code&gt; proves authentication, &lt;code&gt;QUERY_HISTORY&lt;/code&gt; proves who ran what, and &lt;code&gt;ACCESS_HISTORY&lt;/code&gt; proves the exact objects and columns each query read or wrote — the last one is what evidences column-level confidentiality. &lt;code&gt;GRANTS_TO_USERS&lt;/code&gt; and &lt;code&gt;GRANTS_TO_ROLES&lt;/code&gt; snapshot the access model itself for access reviews. BigQuery exposes the same idea through Cloud Audit Logs and &lt;code&gt;INFORMATION_SCHEMA.JOBS&lt;/code&gt;, and Redshift through its &lt;code&gt;STL&lt;/code&gt;/&lt;code&gt;SVL&lt;/code&gt; system tables plus CloudTrail; in every case you export the logs before they age out of retention.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you automate SOC 2 evidence collection?
&lt;/h3&gt;

&lt;p&gt;You build scheduled pipelines that snapshot each control's evidence into an append-only store: nightly grant snapshots, log exports before retention lapses, and the control-test queries (orphaned grants, PII reads, change-control exceptions) run and stored with a timestamp. Each run records a &lt;code&gt;PASS&lt;/code&gt;/&lt;code&gt;FAIL&lt;/code&gt; plus the underlying evidence, feeding a control-monitoring dashboard that alerts the moment a control flips to failing. The audit then becomes a query over the evidence store rather than a quarterly screenshot scramble.&lt;/p&gt;

&lt;h3&gt;
  
  
  What makes an audit table "immutable" for SOC 2?
&lt;/h3&gt;

&lt;p&gt;Three properties. It is append-only — the evidence role has &lt;code&gt;INSERT&lt;/code&gt; and &lt;code&gt;SELECT&lt;/code&gt; but never &lt;code&gt;UPDATE&lt;/code&gt; or &lt;code&gt;DELETE&lt;/code&gt;, so records cannot be quietly changed. It is tamper-evident through hash-chaining, where each row hashes its contents plus the previous row's hash, so altering any historical row breaks the chain and is detectable. And it is written to write-once-read-many storage (S3 Object Lock, GCS retention lock) so even an administrator cannot delete it before its retention period expires.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every SOC 2 control above, from the least-privilege role hierarchy and the orphaned-grant anti-join to the column-level access-history query and the hash-chained evidence ledger, maps to a hands-on practice room where you write the query against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you prove this control operated over the whole audit period?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/access-control" rel="noopener noreferrer"&gt;Practice access-control problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/sla-monitoring" rel="noopener noreferrer"&gt;SLA-monitoring drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Marimo: The Reactive Python Notebook for Reproducible Data Work</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Fri, 25 Sep 2026 13:32:39 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/marimo-the-reactive-python-notebook-for-reproducible-data-work-5c7h</link>
      <guid>https://dev.to/gowthampotureddi/marimo-the-reactive-python-notebook-for-reproducible-data-work-5c7h</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;marimo reactive python notebook&lt;/code&gt;&lt;/strong&gt; is an open-source Python notebook that fixes the one thing every data person has silently tolerated for a decade: a notebook whose displayed output does not match the code you can see. In a classic notebook you can run cell 5, delete cell 3, edit cell 1, and never rerun the cells in between — so the variables in memory are the residue of a run history nobody recorded. Marimo makes that class of bug impossible by treating your notebook as a &lt;strong&gt;dataflow graph&lt;/strong&gt;: it reads which variables each cell defines and references, wires the cells into a directed acyclic graph, and when you change one cell it automatically reruns every cell that depends on it. There is no stale state to reason about, because there is no state that the code on screen cannot reproduce.&lt;/p&gt;

&lt;p&gt;The second idea is just as consequential and follows from the first: a Marimo notebook is stored as a plain &lt;code&gt;.py&lt;/code&gt; file, not a JSON blob. That means it diffs cleanly in git, imports like any module, runs as a script with &lt;code&gt;python notebook.py&lt;/code&gt;, and serves as an interactive web app with &lt;code&gt;marimo run&lt;/code&gt; — the same file, three faces. This guide walks the four things an interviewer will actually probe — the reactive dataflow DAG, the pure-Python reproducibility model, UI elements bound to variables, and SQL cells backed by DuckDB — and pairs each with a Solution-Tail interview answer: code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F68amxorfjz7vqeck64gy.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F68amxorfjz7vqeck64gy.jpeg" alt="PipeCode blog header for Marimo — bold white headline 'Marimo: the reactive notebook' with subtitle 'Dataflow DAG · Pure Python · Reproducible' and a stylised cell-graph scene on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;data-analysis practice library →&lt;/a&gt;, rehearse frame-shaping on the &lt;a href="https://pipecode.ai/explore/practice/topic/dataframe-basics" rel="noopener noreferrer"&gt;dataframe-basics practice set →&lt;/a&gt;, and design the end-to-end flow on the &lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;pipelines practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why Marimo changes the Python notebook in 2026&lt;/li&gt;
&lt;li&gt;The reactive dataflow DAG&lt;/li&gt;
&lt;li&gt;Pure-Python notebooks &amp;amp; reproducibility&lt;/li&gt;
&lt;li&gt;Interactive UI elements bound to variables&lt;/li&gt;
&lt;li&gt;SQL cells &amp;amp; DuckDB&lt;/li&gt;
&lt;li&gt;Cheat sheet — Marimo recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why Marimo changes the Python notebook in 2026
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Marimo is reactive, not top-to-bottom — that one fact removes hidden state and makes a notebook reproducible
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;Marimo derives the execution order from your code's dependencies, not from the order you happened to click cells, so the notebook you see is always the notebook that ran&lt;/strong&gt;. Everything that makes Marimo attractive to a data team follows from that. There is no "run all from the top and pray" ritual, no &lt;code&gt;execution_count&lt;/code&gt; that reads &lt;code&gt;[47]&lt;/code&gt; next to &lt;code&gt;[2]&lt;/code&gt;, no variable that survives from a cell you already deleted. The graph is the source of truth, and the graph is recomputed from the code every time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The problem Marimo is fixing — hidden state.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Out-of-order execution.&lt;/strong&gt; In a classic notebook, cells can run in any order and the interpreter keeps whatever state that produced. The displayed outputs can reflect code that no longer exists on screen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stale variables.&lt;/strong&gt; Delete the cell that defined &lt;code&gt;df&lt;/code&gt;, and &lt;code&gt;df&lt;/code&gt; lingers in memory. Every downstream cell keeps "working" against a ghost until you restart the kernel and discover it was broken all along.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The reproducibility tax.&lt;/strong&gt; Because none of this is safe, the community habit is "Restart &amp;amp; Run All" before trusting a result — a manual discipline that people forget exactly when it matters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How Marimo removes it.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Static analysis builds a DAG.&lt;/strong&gt; Marimo parses each cell to see which global names it &lt;em&gt;defines&lt;/em&gt; and which it &lt;em&gt;references&lt;/em&gt;, then draws an edge from definer to referencer. That graph is a DAG (cycles are rejected).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Change one cell, dependents rerun.&lt;/strong&gt; Edit a cell and Marimo automatically reruns its downstream cells; delete a cell and Marimo removes its variables and invalidates the cells that used them. State on screen and state in memory can never drift.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic order.&lt;/strong&gt; Execution follows a topological sort of the DAG, so two people (or CI) running the same notebook get the same order and the same result.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you say &lt;strong&gt;"Marimo is reactive — the DAG decides execution order"&lt;/strong&gt; in the first sentence? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you frame the core win as &lt;strong&gt;"no hidden state, so the notebook is reproducible by construction"&lt;/strong&gt;? — required framing.&lt;/li&gt;
&lt;li&gt;Do you mention that a notebook is &lt;strong&gt;a pure &lt;code&gt;.py&lt;/code&gt; file, so it diffs and imports&lt;/strong&gt;? — the practical hook.&lt;/li&gt;
&lt;li&gt;Do you know that a UI element becomes reactive &lt;strong&gt;just by reading its &lt;code&gt;.value&lt;/code&gt; in another cell&lt;/strong&gt;, with no callbacks? — the "you actually used it" tell.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — the bug Marimo makes impossible
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The clearest way to feel the difference is to picture the one sequence that silently corrupts a Jupyter notebook and watch Marimo refuse it. You define &lt;code&gt;x&lt;/code&gt;, define &lt;code&gt;y = x + 1&lt;/code&gt;, then go back and edit &lt;code&gt;x&lt;/code&gt;. In Jupyter, &lt;code&gt;y&lt;/code&gt; is now stale until you remember to rerun it. In Marimo, editing &lt;code&gt;x&lt;/code&gt; reruns &lt;code&gt;y&lt;/code&gt; for you — there is no window in which the two disagree.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Two cells: cell A holds &lt;code&gt;x = 1&lt;/code&gt;, cell B holds &lt;code&gt;y = x + 1&lt;/code&gt; and displays &lt;code&gt;y&lt;/code&gt;. You change cell A to &lt;code&gt;x = 10&lt;/code&gt;. What does &lt;code&gt;y&lt;/code&gt; show in Jupyter versus Marimo?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;cell&lt;/th&gt;
&lt;th&gt;code&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;x = 10&lt;/code&gt; (edited from &lt;code&gt;1&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;y = x + 1&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;          &lt;span class="c1"&gt;# cell A
&lt;/span&gt;
&lt;span class="n"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;       &lt;span class="c1"&gt;# cell B (a separate cell)
&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; In a classic notebook, editing cell A does nothing to cell B until you manually rerun B; &lt;code&gt;y&lt;/code&gt; keeps showing &lt;code&gt;2&lt;/code&gt; while &lt;code&gt;x&lt;/code&gt; is already &lt;code&gt;10&lt;/code&gt;. In Marimo, the moment cell A is edited, the runtime sees that cell B &lt;em&gt;references&lt;/em&gt; &lt;code&gt;x&lt;/code&gt;, marks B as a descendant of A, and reruns B — so &lt;code&gt;y&lt;/code&gt; becomes &lt;code&gt;11&lt;/code&gt; with no action from you. The displayed value can never lag the code because the runtime, not the human, owns the order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;environment&lt;/th&gt;
&lt;th&gt;value of &lt;code&gt;y&lt;/code&gt; after editing &lt;code&gt;x&lt;/code&gt; to 10&lt;/th&gt;
&lt;th&gt;stale?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Classic notebook&lt;/td&gt;
&lt;td&gt;2 (until you rerun B)&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Marimo&lt;/td&gt;
&lt;td&gt;11 (reran automatically)&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If editing one cell can leave another cell's output wrong until you remember to rerun it, you are carrying hidden state — Marimo removes the "until you remember" entirely.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. The reactive dataflow DAG
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Cells are functions, references are edges — the DAG is derived from your code, and one variable lives in exactly one cell
&lt;/h3&gt;

&lt;p&gt;Marimo has one central idea you must be able to explain: &lt;strong&gt;the notebook is a graph whose nodes are cells and whose edges are variable dependencies&lt;/strong&gt;, and Marimo builds it by static analysis, never by running your code speculatively. An interviewer who asks "how does Marimo know what to rerun?" wants this graph, the single-definition rule that keeps it a clean DAG, and the minimal-recompute behaviour that falls out of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How the graph is built.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Defs and refs.&lt;/strong&gt; Each cell is analysed for the global names it &lt;strong&gt;defines&lt;/strong&gt; (assignments, &lt;code&gt;def&lt;/code&gt;, &lt;code&gt;import&lt;/code&gt;, &lt;code&gt;class&lt;/code&gt;) and the names it &lt;strong&gt;references&lt;/strong&gt;. A cell that references &lt;code&gt;df&lt;/code&gt; depends on whichever cell defines &lt;code&gt;df&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edges point from producer to consumer.&lt;/strong&gt; If cell 1 defines &lt;code&gt;df&lt;/code&gt; and cell 3 uses &lt;code&gt;df&lt;/code&gt;, there is an edge 1 → 3. Cell 3 is a &lt;em&gt;descendant&lt;/em&gt; of cell 1.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It must be acyclic.&lt;/strong&gt; Two cells cannot mutually depend on each other's variables; Marimo rejects cycles because a cycle has no valid execution order.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The single-definition rule.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One variable, one cell.&lt;/strong&gt; A given global name may be defined in exactly one cell. Redefining &lt;code&gt;df&lt;/code&gt; in a second cell is an error, not a silent overwrite — this is what guarantees the graph is well-defined.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Underscore for locals.&lt;/strong&gt; A name prefixed with &lt;code&gt;_&lt;/code&gt; (e.g. &lt;code&gt;_tmp&lt;/code&gt;) is cell-local and invisible to the graph, so you can reuse throwaway names freely without touching the DAG.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why the rule exists.&lt;/strong&gt; If two cells could both define &lt;code&gt;df&lt;/code&gt;, "which one wins" would depend on run order — exactly the hidden state Marimo abolishes. Forbidding it keeps the DAG unambiguous.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What the runtime does with the DAG.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Minimal recompute.&lt;/strong&gt; When a cell changes, Marimo reruns that cell and only its transitive descendants — not the whole notebook. Unrelated branches are left alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delete = invalidate.&lt;/strong&gt; Deleting a cell removes its variables from memory and reruns the cells that referenced them, so nothing keeps working against a ghost variable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Topological execution.&lt;/strong&gt; On a fresh run, cells execute in dependency order regardless of their visual position on the page, which is why you can arrange cells however reads best.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgpoqffktmhzi8iaj66nr.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgpoqffktmhzi8iaj66nr.jpeg" alt="Iconographic Marimo dataflow diagram — notebook cells parsed into defs and refs, a directed acyclic graph linking them, the single-definition rule, and a changed cell rerunning only its downstream dependents." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — one edit, only the dependents rerun
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The payoff of the DAG is surgical recomputation. Build a tiny four-cell notebook where a raw frame feeds a filter, the filter feeds a chart, and a wholly unrelated cell computes a constant. Change the raw frame and watch exactly two cells rerun while the unrelated cell sits still.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Given cells that define &lt;code&gt;raw&lt;/code&gt;, &lt;code&gt;clean&lt;/code&gt; (depends on &lt;code&gt;raw&lt;/code&gt;), &lt;code&gt;chart&lt;/code&gt; (depends on &lt;code&gt;clean&lt;/code&gt;), and &lt;code&gt;note&lt;/code&gt; (independent), which cells rerun when you edit &lt;code&gt;raw&lt;/code&gt;?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;cell&lt;/th&gt;
&lt;th&gt;defines&lt;/th&gt;
&lt;th&gt;references&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;raw&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;&lt;code&gt;clean&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;raw&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;chart&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;clean&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;&lt;code&gt;note&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;load_orders&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                      &lt;span class="c1"&gt;# cell 1
&lt;/span&gt;
&lt;span class="n"&gt;clean&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;           &lt;span class="c1"&gt;# cell 2
&lt;/span&gt;
&lt;span class="n"&gt;chart&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;clean&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;plot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;day&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# cell 3
&lt;/span&gt;
&lt;span class="n"&gt;note&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dashboard v2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;                    &lt;span class="c1"&gt;# cell 4 (independent of the chain)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Marimo reads the four cells and draws edges 1 → 2 (both touch &lt;code&gt;raw&lt;/code&gt;), 2 → 3 (both touch &lt;code&gt;clean&lt;/code&gt;), and leaves cell 4 with no edges. When you edit cell 1, the runtime walks the descendants of node 1: cell 2 reruns because it references &lt;code&gt;raw&lt;/code&gt;, then cell 3 reruns because it references &lt;code&gt;clean&lt;/code&gt;. Cell 4 never runs because &lt;code&gt;note&lt;/code&gt; depends on nothing that changed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;edited cell&lt;/th&gt;
&lt;th&gt;cells that rerun&lt;/th&gt;
&lt;th&gt;cells left untouched&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 (&lt;code&gt;raw&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;2 (&lt;code&gt;clean&lt;/code&gt;), 3 (&lt;code&gt;chart&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;4 (&lt;code&gt;note&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Marimo reruns the &lt;em&gt;transitive descendants&lt;/em&gt; of what you changed — nothing upstream, nothing on a sibling branch — so an expensive independent cell is never recomputed by an unrelated edit.&lt;/p&gt;

&lt;h3&gt;
  
  
  Marimo interview question on the dependency graph
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; An interviewer shows you a notebook where cell 1 defines &lt;code&gt;df&lt;/code&gt;, cell 2 also tries to define &lt;code&gt;df = df.dropna()&lt;/code&gt;, and cell 3 reads &lt;code&gt;df&lt;/code&gt;. They ask why Marimo rejects this and how you would rewrite it so the intent (drop nulls, then use the clean frame) works reactively. Show the fix.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using distinct variables to keep the DAG acyclic
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;load_orders&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                    &lt;span class="c1"&gt;# cell 1 — the raw frame, defined once
&lt;/span&gt;
&lt;span class="n"&gt;df_clean&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dropna&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                &lt;span class="c1"&gt;# cell 2 — new name (single-definition rule satisfied)
&lt;/span&gt;
&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df_clean&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;groupby&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;   &lt;span class="c1"&gt;# cell 3 — consume the cleaned frame
&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;cell&lt;/th&gt;
&lt;th&gt;attempted defs&lt;/th&gt;
&lt;th&gt;valid under single-definition rule?&lt;/th&gt;
&lt;th&gt;edge added&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;df&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 (original)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;df&lt;/code&gt; again&lt;/td&gt;
&lt;td&gt;no — redefinition error&lt;/td&gt;
&lt;td&gt;rejected&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 (fixed)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;df_clean&lt;/code&gt; (refs &lt;code&gt;df&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;1 → 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;result&lt;/code&gt; (refs &lt;code&gt;df_clean&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;2 → 3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The original cell 2 redefines &lt;code&gt;df&lt;/code&gt;, which Marimo flags as a &lt;strong&gt;multiple-definition&lt;/strong&gt; error — two cells owning one name would make execution order ambiguous.&lt;/li&gt;
&lt;li&gt;Renaming the cleaned frame to &lt;code&gt;df_clean&lt;/code&gt; gives each name a single owner, so the graph stays a valid DAG.&lt;/li&gt;
&lt;li&gt;Now editing cell 1 reruns cell 2 (recomputes &lt;code&gt;df_clean&lt;/code&gt;) and then cell 3 (recomputes &lt;code&gt;result&lt;/code&gt;) automatically.&lt;/li&gt;
&lt;li&gt;Because &lt;code&gt;df&lt;/code&gt; and &lt;code&gt;df_clean&lt;/code&gt; are distinct nodes, you can inspect both the raw and cleaned frames at any time without one clobbering the other.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;frame&lt;/th&gt;
&lt;th&gt;shape&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;cell 1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;df&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1000 rows (with nulls)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cell 2&lt;/td&gt;
&lt;td&gt;&lt;code&gt;df_clean&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;960 rows (nulls dropped)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cell 3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;result&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;one sum per customer&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Single-definition rule&lt;/strong&gt;&lt;/strong&gt; — one global name per cell makes the DAG unambiguous; Marimo can always answer "which cell owns &lt;code&gt;df&lt;/code&gt;?" with exactly one node.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Distinct names as nodes&lt;/strong&gt;&lt;/strong&gt; — turning the mutation into a new variable &lt;code&gt;df_clean&lt;/code&gt; adds a node and an edge instead of a hidden overwrite, so the transformation becomes a visible, rerunnable step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Acyclic guarantee&lt;/strong&gt;&lt;/strong&gt; — because no cell redefines an upstream name, the graph has a topological order and every edit has a well-defined set of cells to rerun.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Reactive propagation&lt;/strong&gt;&lt;/strong&gt; — a change to &lt;code&gt;df&lt;/code&gt; flows deterministically to &lt;code&gt;df_clean&lt;/code&gt; and then &lt;code&gt;result&lt;/code&gt;, with the runtime, not the analyst, choosing the order.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — building the graph is O(cells × names) of cheap static parsing; a rerun touches only O(descendants), not the whole notebook.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Analysis&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — data-analysis&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Exploratory data-analysis notebook problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;DataFrames&lt;/span&gt;
&lt;span&gt;Topic — dataframe-basics&lt;/span&gt;
&lt;strong&gt;DataFrame define-and-transform problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/dataframe-basics" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. Pure-Python notebooks &amp;amp; reproducibility
&lt;/h2&gt;
&lt;h3&gt;
  
  
  The notebook is a &lt;code&gt;.py&lt;/code&gt; file, so it diffs, imports, and runs three ways — reproducibility is a property of the format, not a discipline
&lt;/h3&gt;

&lt;p&gt;The feature that sells Marimo to an engineer is that &lt;strong&gt;a notebook is stored as an ordinary Python module, not a JSON document&lt;/strong&gt;. A &lt;code&gt;.ipynb&lt;/code&gt; file interleaves source, base64 outputs, and execution metadata into JSON that produces unreadable diffs and merge conflicts; a Marimo file is code you can read, review, and run without a kernel. Reproducibility stops being a habit you enforce and becomes a property of the artifact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the file actually is.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A valid module.&lt;/strong&gt; The notebook is a &lt;code&gt;.py&lt;/code&gt; file defining an &lt;code&gt;app = marimo.App()&lt;/code&gt; with each cell as a small decorated function that returns its defined names. It executes as plain Python.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clean git diffs.&lt;/strong&gt; Because it is code, a one-line change is a one-line diff — no serialized output blobs, no &lt;code&gt;execution_count&lt;/code&gt; churn, no &lt;code&gt;outputs: []&lt;/code&gt; noise. Code review works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No hidden outputs on disk.&lt;/strong&gt; Cell outputs are not baked into the file, so you cannot accidentally commit a stale chart or a leaked credential printed three runs ago.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Deterministic execution, guaranteed.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Order from the DAG, not the page.&lt;/strong&gt; On any run — yours, a colleague's, CI's — cells execute in topological order of the dependency graph, so the result does not depend on click history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No out-of-order bugs.&lt;/strong&gt; The one class of notebook bug that "works on my machine" comes from run order; Marimo eliminates it because run order is derived, not remembered.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reproducible by construction.&lt;/strong&gt; Given the same inputs and the same file, the same outputs follow — which is exactly what a data pipeline or a graded assignment needs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Three run modes from one file.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;marimo edit notebook.py&lt;/code&gt;&lt;/strong&gt; opens the reactive editor for development.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;marimo run notebook.py&lt;/code&gt;&lt;/strong&gt; serves it as a read-only interactive web app — the code is hidden and only the UI and outputs show, so a notebook becomes a dashboard with no rewrite.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;python notebook.py&lt;/code&gt;&lt;/strong&gt; runs it as a script (great for a cron job or an Airflow task); because it is importable, you can also &lt;code&gt;from notebook import result&lt;/code&gt; and unit-test cells with pytest.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fetbjbutv21l33h8hde50.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fetbjbutv21l33h8hde50.jpeg" alt="Iconographic Marimo reproducibility diagram — a notebook stored as a pure .py module with a clean git diff, a deterministic execution order from the DAG, and three run modes: edit, app, and script." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — the same notebook as a script
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The most convincing demonstration of the pure-Python format is running a notebook headless. Because the file is a real module with a &lt;code&gt;__main__&lt;/code&gt; guard, &lt;code&gt;python notebook.py&lt;/code&gt; executes every cell in dependency order and any top-level output (or a &lt;code&gt;mo.cli_output&lt;/code&gt;) is produced without ever opening a browser — the notebook and the batch job are the same code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You have &lt;code&gt;analysis.py&lt;/code&gt;, authored in Marimo, that loads data and computes a summary. Show that it runs unchanged as a batch script and yields the same summary as the editor.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; A notebook whose last cell defines &lt;code&gt;summary&lt;/code&gt; from an upstream &lt;code&gt;orders&lt;/code&gt; frame.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;marimo&lt;/span&gt;    &lt;span class="c1"&gt;# file: analysis.py — authored in Marimo, stored as pure Python
&lt;/span&gt;&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;marimo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;App&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="nd"&gt;@app.cell&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;polars&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pl&lt;/span&gt;
    &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_parquet&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;orders.parquet&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pl&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nd"&gt;@app.cell&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;agg&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;total&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sum&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;,)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The file is legal Python: importing it or running it triggers &lt;code&gt;app.run()&lt;/code&gt;, which executes the cells in DAG order — the &lt;code&gt;orders&lt;/code&gt; cell first because the &lt;code&gt;summary&lt;/code&gt; cell references &lt;code&gt;orders&lt;/code&gt;. Nothing in the file depends on a browser or a saved kernel, so the batch run reproduces the editor run exactly. Because each cell returns its defined names, another module can &lt;code&gt;from analysis import summary&lt;/code&gt; and assert on it in a test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;invocation&lt;/th&gt;
&lt;th&gt;executes&lt;/th&gt;
&lt;th&gt;produces&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;marimo edit analysis.py&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cells in DAG order (interactive)&lt;/td&gt;
&lt;td&gt;live &lt;code&gt;summary&lt;/code&gt; in the editor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;python analysis.py&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;same cells, same order (headless)&lt;/td&gt;
&lt;td&gt;identical &lt;code&gt;summary&lt;/code&gt;, no browser&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If your notebook must also be a scheduled job or pass code review, choose a format that is already Python — reproducibility you get for free beats reproducibility you have to remember.&lt;/p&gt;

&lt;h3&gt;
  
  
  Marimo interview question on reproducibility
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A teammate reports that a &lt;code&gt;.ipynb&lt;/code&gt; gives different numbers depending on who runs it, and blames "the data." You suspect out-of-order execution. Explain how moving to Marimo removes that failure mode, and what specifically guarantees a deterministic result.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using DAG-ordered execution instead of click order
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;rate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;1.08&lt;/span&gt;                      &lt;span class="c1"&gt;# cell A — defines rate
&lt;/span&gt;
&lt;span class="n"&gt;adjusted&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;base_amounts&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;rate&lt;/span&gt;   &lt;span class="c1"&gt;# cell B — uses rate to build adjusted
&lt;/span&gt;
&lt;span class="n"&gt;report&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;adjusted&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;     &lt;span class="c1"&gt;# cell C — uses adjusted to build report
&lt;/span&gt;&lt;span class="n"&gt;report&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;environment&lt;/th&gt;
&lt;th&gt;how order is chosen&lt;/th&gt;
&lt;th&gt;result if cells were clicked A, C, B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Classic notebook&lt;/td&gt;
&lt;td&gt;human click order&lt;/td&gt;
&lt;td&gt;C sees old/undefined &lt;code&gt;adjusted&lt;/code&gt; → wrong or error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Marimo&lt;/td&gt;
&lt;td&gt;topological sort of the DAG&lt;/td&gt;
&lt;td&gt;always A → B → C → correct &lt;code&gt;report&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The three cells form a chain A → B → C because B references &lt;code&gt;rate&lt;/code&gt; and C references &lt;code&gt;adjusted&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;In a classic notebook, running them A, C, B produces a &lt;code&gt;report&lt;/code&gt; built from a stale or missing &lt;code&gt;adjusted&lt;/code&gt;, and the number depends on run history — not the data.&lt;/li&gt;
&lt;li&gt;Marimo ignores the order you interact with cells and executes in topological order every time, so A runs before B runs before C, deterministically.&lt;/li&gt;
&lt;li&gt;Because the file is pure Python with no stored outputs, a second person cloning the repo and running it gets byte-identical execution order.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;run by&lt;/th&gt;
&lt;th&gt;execution order&lt;/th&gt;
&lt;th&gt;&lt;code&gt;report&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;author (Marimo)&lt;/td&gt;
&lt;td&gt;A → B → C&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;teammate (Marimo)&lt;/td&gt;
&lt;td&gt;A → B → C&lt;/td&gt;
&lt;td&gt;identical&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Topological order&lt;/strong&gt;&lt;/strong&gt; — deriving execution order from dependencies means the answer cannot depend on click history, which is the root cause of "works on my machine" notebooks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Pure-Python format&lt;/strong&gt;&lt;/strong&gt; — no serialized outputs and no &lt;code&gt;execution_count&lt;/code&gt; means the file that runs is exactly the file in git; there is nothing hidden to diverge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Importable module&lt;/strong&gt;&lt;/strong&gt; — because cells return their names, the notebook can be imported and unit-tested, turning "trust me" into an assertion in CI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Determinism as a default&lt;/strong&gt;&lt;/strong&gt; — reproducibility is guaranteed by the runtime rather than by a "Restart &amp;amp; Run All" ritual people forget.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the topological sort is O(cells + edges) once per run; the reproducibility it buys is unbounded in engineering time saved.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Pipelines&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — pipelines&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Reproducible pipeline-ordering problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Analysis&lt;/span&gt;
&lt;span&gt;Topic — data-analysis&lt;/span&gt;
&lt;strong&gt;End-to-end analysis reproducibility problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Interactive UI elements bound to variables
&lt;/h2&gt;
&lt;h3&gt;
  
  
  A widget is a Python object; reading its &lt;code&gt;.value&lt;/code&gt; in another cell wires reactivity — no callbacks, no state juggling
&lt;/h3&gt;

&lt;p&gt;Marimo's interactivity is the reactive DAG applied to human input. A UI element — &lt;code&gt;mo.ui.slider&lt;/code&gt;, &lt;code&gt;mo.ui.dropdown&lt;/code&gt;, &lt;code&gt;mo.ui.text&lt;/code&gt;, &lt;code&gt;mo.ui.table&lt;/code&gt; — is just a Python object with a live &lt;code&gt;.value&lt;/code&gt;. The instant you &lt;strong&gt;read that &lt;code&gt;.value&lt;/code&gt; in a different cell&lt;/strong&gt;, that cell becomes a dependent of the widget, so moving the slider reruns exactly the cells that consume it. You never register a callback, never mutate shared state, never wire an event handler — the graph does it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The binding model.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The element is a value.&lt;/strong&gt; &lt;code&gt;slider = mo.ui.slider(1, 100, value=10)&lt;/code&gt; creates an object; displaying it renders the control, and &lt;code&gt;slider.value&lt;/code&gt; reads the current position as a plain Python number.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Define and display in one cell, read in another.&lt;/strong&gt; The reactive contract is: the widget is defined and shown in its own cell, and &lt;em&gt;other&lt;/em&gt; cells read &lt;code&gt;slider.value&lt;/code&gt;. Reading it elsewhere is what draws the dependency edge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interaction = a cell edit.&lt;/strong&gt; Dragging the slider is, to the runtime, equivalent to changing the cell that owns &lt;code&gt;slider&lt;/code&gt; — so Marimo reruns the widget's descendants and nothing else.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why there are no callbacks.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No &lt;code&gt;on_change&lt;/code&gt; handlers.&lt;/strong&gt; Traditional widget libraries make you attach a function that fires on change and mutate globals; that is imperative state management and it is where dashboards rot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reactivity replaces events.&lt;/strong&gt; Because the dependent cell already declares what it needs (&lt;code&gt;slider.value&lt;/code&gt;), Marimo knows what to rerun without you describing &lt;em&gt;when&lt;/em&gt;. Declarative beats imperative.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A dashboard for free.&lt;/strong&gt; Pair the widgets and their dependent charts, run &lt;code&gt;marimo run&lt;/code&gt;, and the notebook is an interactive app — same file, no framework.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The control knobs.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;mo.stop(predicate, output)&lt;/code&gt;&lt;/strong&gt; short-circuits a cell and its descendants when a guard is true (e.g. no file uploaded yet), so an expensive graph does not run on empty input.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;mo.ui.form&lt;/code&gt;&lt;/strong&gt; wraps inputs so the graph reruns on submit rather than on every keystroke — the right tool for costly downstream work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compose with &lt;code&gt;mo.ui.array&lt;/code&gt; / &lt;code&gt;mo.ui.dictionary&lt;/code&gt;&lt;/strong&gt; to build a list or map of elements and read them all through one &lt;code&gt;.value&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnkulicvjq6b3psyn1zgs.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnkulicvjq6b3psyn1zgs.jpeg" alt="Iconographic Marimo UI diagram — a slider and dropdown widget whose .value is read by a downstream cell, so moving the slider reruns only the dependent chart cell, with no callback wiring." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — a slider that filters a frame
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The canonical interactive pattern is a slider that sets a threshold and a downstream cell that filters a dataframe by it. Define the slider in one cell, read &lt;code&gt;slider.value&lt;/code&gt; in the filter cell, and the filtered frame (and any chart built from it) updates the moment you drag — with zero event wiring.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Build a &lt;code&gt;min_amount&lt;/code&gt; slider from 0 to 100 and a cell that shows only the &lt;code&gt;orders&lt;/code&gt; rows whose &lt;code&gt;amount&lt;/code&gt; is at least the slider value. What reruns when you drag it to 50?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_id&lt;/th&gt;
&lt;th&gt;amount&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;marimo&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;

&lt;span class="n"&gt;min_amount&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ui&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;label&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;min amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# cell 1 — define + display
&lt;/span&gt;&lt;span class="n"&gt;min_amount&lt;/span&gt;

&lt;span class="n"&gt;filtered&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;min_amount&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;   &lt;span class="c1"&gt;# cell 2 — reads .value, depends on slider
&lt;/span&gt;&lt;span class="n"&gt;filtered&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Cell 1 creates the slider object and displays the control; &lt;code&gt;min_amount.value&lt;/code&gt; starts at &lt;code&gt;0&lt;/code&gt;. Cell 2 reads &lt;code&gt;min_amount.value&lt;/code&gt;, so Marimo draws an edge from the slider cell to the filter cell. Dragging the slider to &lt;code&gt;50&lt;/code&gt; is treated as a change to cell 1, so Marimo reruns cell 2 (and any chart cell below it) — recomputing &lt;code&gt;filtered&lt;/code&gt; with the new threshold. No &lt;code&gt;on_change&lt;/code&gt;, no global mutation: the dependency edge already said what to do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;slider value&lt;/th&gt;
&lt;th&gt;rows in &lt;code&gt;filtered&lt;/code&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1, 2, 3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;2, 3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;80&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Define-and-display the widget in one cell, read &lt;code&gt;.value&lt;/code&gt; in the cells that consume it — the moment you read the value elsewhere, the widget is wired reactively with no callback.&lt;/p&gt;

&lt;h3&gt;
  
  
  Marimo interview question on interactive reactivity
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You have an expensive model-fit cell that reads a &lt;code&gt;mo.ui.dropdown&lt;/code&gt; of dataset names and a &lt;code&gt;mo.ui.slider&lt;/code&gt; of epochs. The interviewer wants the fit to run only when the user explicitly submits — not on every twitch of the slider — while still being fully reactive. How do you build it?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using mo.ui.form to batch inputs before rerun
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;marimo&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;

&lt;span class="n"&gt;controls&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ui&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dictionary&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;    &lt;span class="c1"&gt;# cell 1 — form reruns on SUBMIT, not per keystroke
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dataset&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ui&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dropdown&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;c&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;epochs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;  &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ui&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;form&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;controls&lt;/span&gt;

&lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;controls&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;md&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Set options and press **Submit**.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;  &lt;span class="c1"&gt;# cell 2 — guard
&lt;/span&gt;&lt;span class="n"&gt;params&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;controls&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;                       &lt;span class="c1"&gt;# only set after submit
&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;expensive_fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dataset&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;epochs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;user action&lt;/th&gt;
&lt;th&gt;&lt;code&gt;controls.value&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;cell 2 behaviour&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;dragging slider (no submit)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;None&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;mo.stop&lt;/code&gt; halts — no fit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;press Submit&lt;/td&gt;
&lt;td&gt;&lt;code&gt;{"dataset": "b", "epochs": 30}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;guard passes → &lt;code&gt;expensive_fit&lt;/code&gt; runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;drag again (no submit)&lt;/td&gt;
&lt;td&gt;still last submitted&lt;/td&gt;
&lt;td&gt;fit does not rerun&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Wrapping the inputs in &lt;code&gt;.form()&lt;/code&gt; means their &lt;code&gt;.value&lt;/code&gt; only updates on &lt;strong&gt;Submit&lt;/strong&gt;, so mid-drag changes do not touch the graph.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;mo.stop(controls.value is None, ...)&lt;/code&gt; short-circuits cell 2 (and its descendants) until the first submit, so the expensive fit never runs on empty input.&lt;/li&gt;
&lt;li&gt;After submit, &lt;code&gt;controls.value&lt;/code&gt; is a dict of the chosen options; cell 2 reads it and runs &lt;code&gt;expensive_fit&lt;/code&gt; exactly once.&lt;/li&gt;
&lt;li&gt;Because the form is still a normal reactive object, a later submit reruns only cell 2 downstream — the reactivity is intact, just batched.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;phase&lt;/th&gt;
&lt;th&gt;fit runs?&lt;/th&gt;
&lt;th&gt;why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;before first submit&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;mo.stop&lt;/code&gt; guard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;on submit&lt;/td&gt;
&lt;td&gt;yes (once)&lt;/td&gt;
&lt;td&gt;form value materialized&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;twiddling after submit&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;value unchanged until next submit&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Widget as value&lt;/strong&gt;&lt;/strong&gt; — every input is a Python object whose &lt;code&gt;.value&lt;/code&gt; participates in the DAG, so no event handlers are needed to connect input to computation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Form batching&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;.form()&lt;/code&gt; defers the value update to submit, converting "rerun on every keystroke" into "rerun on intent," which is what expensive cells need.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;mo.stop guard&lt;/strong&gt;&lt;/strong&gt; — short-circuiting the cell and its descendants keeps a costly graph dormant until inputs are valid, the reactive analogue of an early return.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Declarative reactivity&lt;/strong&gt;&lt;/strong&gt; — the dependent cell states what it reads, so Marimo reruns the right cells without you specifying when, eliminating callback spaghetti.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the fit is O(fit) and now runs only on submit; the reactive bookkeeping is O(descendants of the form), independent of how much the user fiddles.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;DataFrames&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — dataframe-basics&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Interactive filter-and-aggregate problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/dataframe-basics" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Analysis&lt;/span&gt;
&lt;span&gt;Topic — data-analysis&lt;/span&gt;
&lt;strong&gt;Parameterized dashboard analysis problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. SQL cells &amp;amp; DuckDB
&lt;/h2&gt;
&lt;h3&gt;
  
  
  &lt;code&gt;mo.sql&lt;/code&gt; runs SQL over your Python dataframes with embedded DuckDB and hands back a dataframe — SQL and Python share one namespace
&lt;/h3&gt;

&lt;p&gt;Marimo lets you write SQL as a first-class cell, and the engine underneath is embedded &lt;strong&gt;DuckDB&lt;/strong&gt;. The magic is that a SQL cell can reference your in-memory Python dataframes &lt;em&gt;by their variable name&lt;/em&gt; — no load step, no connection string — and it returns the query result as a dataframe that is itself a reactive variable. SQL feeds Python, which feeds more SQL, all inside the same DAG.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How a SQL cell works.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;mo.sql(...)&lt;/code&gt; is the primitive.&lt;/strong&gt; In the editor a SQL cell is authored as SQL, but it compiles to &lt;code&gt;mo.sql("SELECT ...")&lt;/code&gt;; it executes the query on DuckDB and returns a dataframe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dataframes are tables.&lt;/strong&gt; A Python dataframe named &lt;code&gt;orders&lt;/code&gt; is queryable as &lt;code&gt;FROM orders&lt;/code&gt; directly — DuckDB reads the pandas/polars frame in place, so there is no import or copy into a database.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The result is a reactive variable.&lt;/strong&gt; Name the output (e.g. &lt;code&gt;big_orders&lt;/code&gt;) and it becomes a node in the DAG; downstream Python or SQL cells that reference it rerun when the query changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Parameterizing and mixing.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;f-string interpolation.&lt;/strong&gt; Because the SQL is a Python string, you interpolate values with &lt;code&gt;{...}&lt;/code&gt; — &lt;code&gt;WHERE amount &amp;gt; {min_amount.value}&lt;/code&gt; wires a UI slider straight into a query, and moving the slider reruns the SQL cell.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SQL ⇄ Python round-trips.&lt;/strong&gt; Clean in pandas, aggregate in SQL, chart in Python — each step is a cell, each output is a frame, and the DAG keeps them in sync.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One embedded engine.&lt;/strong&gt; DuckDB runs in-process, so there is no server to provision; it is columnar and vectorized, so group-bys over millions of rows are fast on a laptop.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers probe.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reactivity of results&lt;/strong&gt; — the query output is a normal reactive frame, so a change upstream reruns the SQL and everything after it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No hidden connection state&lt;/strong&gt; — because DuckDB is embedded and the frames are in memory, there is no external database whose state could drift from the notebook.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When SQL beats pandas&lt;/strong&gt; — set-based joins and aggregations read more clearly as SQL; row-wise Python logic stays in Python. Use each where it is strongest.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw3iwostw7m72c5totosw.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw3iwostw7m72c5totosw.jpeg" alt="Iconographic Marimo SQL diagram — a mo.sql cell querying a Python dataframe by name through an embedded DuckDB engine and returning a new dataframe that a downstream Python cell consumes." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — SQL over a Python dataframe
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The clearest demonstration is a SQL cell that aggregates an in-memory dataframe and returns a result you keep working with in Python. No &lt;code&gt;CREATE TABLE&lt;/code&gt;, no &lt;code&gt;INSERT&lt;/code&gt;, no connection — you name the frame in the &lt;code&gt;FROM&lt;/code&gt; clause and DuckDB reads it directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Given a Python dataframe &lt;code&gt;orders(customer, amount)&lt;/code&gt;, write a SQL cell that returns each customer's total spend, and show that the result is a dataframe you can use downstream.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;amount&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;linus&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;marimo&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;

&lt;span class="n"&gt;totals&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sql&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;          &lt;span class="c1"&gt;# SQL cell — returns a dataframe named `totals`
&lt;/span&gt;    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    SELECT customer, SUM(amount) AS total
    FROM orders
    GROUP BY customer
    ORDER BY total DESC
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;top_customer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;iloc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;   &lt;span class="c1"&gt;# downstream Python cell consumes it reactively
&lt;/span&gt;&lt;span class="n"&gt;top_customer&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;mo.sql(...)&lt;/code&gt; hands the query to the embedded DuckDB engine, which resolves &lt;code&gt;FROM orders&lt;/code&gt; against the in-memory Python frame — no load step. The engine groups by &lt;code&gt;customer&lt;/code&gt;, sums &lt;code&gt;amount&lt;/code&gt;, and returns the result as a dataframe bound to &lt;code&gt;totals&lt;/code&gt;. Because &lt;code&gt;totals&lt;/code&gt; is a normal reactive variable, the downstream cell reading &lt;code&gt;totals.iloc[0]&lt;/code&gt; becomes its dependent, so if &lt;code&gt;orders&lt;/code&gt; changes upstream, both the SQL cell and &lt;code&gt;top_customer&lt;/code&gt; rerun.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;linus&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If a step is a join or a group-by, reach for a &lt;code&gt;mo.sql&lt;/code&gt; cell over your dataframe — DuckDB reads the frame in place and the result is just another reactive frame.&lt;/p&gt;

&lt;h3&gt;
  
  
  Marimo interview question on parameterized SQL
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You want a &lt;code&gt;mo.ui.slider&lt;/code&gt; to set a minimum spend, and a SQL cell that returns only customers above that threshold, updating live as the slider moves. Show how the slider value reaches the SQL and why the query reruns.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using f-string interpolation from a UI element into SQL
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;marimo&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;

&lt;span class="n"&gt;min_spend&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ui&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;label&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;min spend&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# cell 1 — threshold slider
&lt;/span&gt;&lt;span class="n"&gt;min_spend&lt;/span&gt;

&lt;span class="n"&gt;big_spenders&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sql&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;          &lt;span class="c1"&gt;# cell 2 — SQL interpolates the slider value
&lt;/span&gt;    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    SELECT customer, SUM(amount) AS total
    FROM orders
    GROUP BY customer
    HAVING SUM(amount) &amp;gt;= &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;min_spend&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
    ORDER BY total DESC
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;big_spenders&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;slider value&lt;/th&gt;
&lt;th&gt;interpolated &lt;code&gt;HAVING&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;rows returned&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;&lt;code&gt;&amp;gt;= 0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;ada (100), linus (15)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;&lt;code&gt;&amp;gt;= 50&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;ada (100)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;&lt;code&gt;&amp;gt;= 100&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;ada (100)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Cell 2 reads &lt;code&gt;min_spend.value&lt;/code&gt; inside the f-string, so Marimo draws an edge from the slider cell to the SQL cell.&lt;/li&gt;
&lt;li&gt;Dragging the slider is treated as a change to cell 1, so Marimo reruns cell 2 — DuckDB re-executes the query with the new &lt;code&gt;HAVING&lt;/code&gt; bound.&lt;/li&gt;
&lt;li&gt;DuckDB resolves &lt;code&gt;FROM orders&lt;/code&gt; against the in-memory frame each run; no reload or connection is involved.&lt;/li&gt;
&lt;li&gt;The result &lt;code&gt;big_spenders&lt;/code&gt; is a reactive frame, so any chart cell below it reruns too — the whole chain from slider to SQL to chart stays in sync.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;slider&lt;/th&gt;
&lt;th&gt;&lt;code&gt;big_spenders&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;ada, linus&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;ada&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Embedded DuckDB&lt;/strong&gt;&lt;/strong&gt; — an in-process columnar engine queries Python frames by name, so there is no server, no connection state, and no copy into a database.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;f-string parameterization&lt;/strong&gt;&lt;/strong&gt; — because SQL is a Python string, a UI element's &lt;code&gt;.value&lt;/code&gt; interpolates directly, wiring a slider into a query through the same DAG that governs everything else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Result is a reactive frame&lt;/strong&gt;&lt;/strong&gt; — naming the query output makes it a node, so SQL can feed Python which can feed more SQL, all kept consistent by the runtime.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Live re-execution&lt;/strong&gt;&lt;/strong&gt; — moving the slider reruns only the SQL cell and its descendants, giving an interactive query with no callback or refresh button.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — each rerun is O(query) on DuckDB's vectorized engine over the in-memory frame; the reactive overhead is O(descendants of the slider), not the whole notebook.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Analysis&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — data-analysis&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;SQL-over-dataframe analysis problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Pipelines&lt;/span&gt;
&lt;span&gt;Topic — pipelines&lt;/span&gt;
&lt;strong&gt;SQL-and-Python pipeline problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — Marimo recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Minimal reactive cell pair.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;21&lt;/span&gt;          &lt;span class="c1"&gt;# cell 1
&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;       &lt;span class="c1"&gt;# cell 2 — reruns automatically when x changes
&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;UI element bound to a variable.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;marimo&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;
&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ui&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# define + display in one cell
&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;
&lt;span class="n"&gt;squared&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;               &lt;span class="c1"&gt;# in another cell: reading .value wires reactivity
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;SQL cell over a dataframe (embedded DuckDB).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sql&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT customer, SUM(amount) AS total FROM orders GROUP BY customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Guard an expensive cell.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;uploaded&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;md&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Upload a file to continue.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;expensive_fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;uploaded&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Run three ways from one file.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;marimo edit notebook.py     &lt;span class="c"&gt;# reactive editor&lt;/span&gt;
marimo run  notebook.py     &lt;span class="c"&gt;# read-only interactive app (code hidden)&lt;/span&gt;
python      notebook.py     &lt;span class="c"&gt;# headless script (cron / Airflow)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Sandbox with inline dependencies (reproducible env).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;marimo edit &lt;span class="nt"&gt;--sandbox&lt;/span&gt; notebook.py   &lt;span class="c"&gt;# deps pinned in the file via uv (PEP 723)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Choosing where logic lives.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Put it in&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Join / group-by / filter over a frame&lt;/td&gt;
&lt;td&gt;a &lt;code&gt;mo.sql&lt;/code&gt; cell (DuckDB)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Row-wise Python / model fit&lt;/td&gt;
&lt;td&gt;a Python cell&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A value a human should tune&lt;/td&gt;
&lt;td&gt;a &lt;code&gt;mo.ui&lt;/code&gt; element read via &lt;code&gt;.value&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reuse across notebooks / tests&lt;/td&gt;
&lt;td&gt;a plain function imported into a cell&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is Marimo (the reactive Python notebook)?
&lt;/h3&gt;

&lt;p&gt;Marimo is an open-source Python notebook that runs &lt;strong&gt;reactively&lt;/strong&gt;: it builds a dependency graph from your code and, when you change a cell, automatically reruns every cell that depends on it. There is no hidden state, because the runtime — not your click history — decides execution order. Notebooks are stored as pure &lt;code&gt;.py&lt;/code&gt; files, so they diff in git, import as modules, and run as scripts or apps.&lt;/p&gt;

&lt;h3&gt;
  
  
  How is Marimo different from Jupyter?
&lt;/h3&gt;

&lt;p&gt;Jupyter executes cells in whatever order you click them and stores the notebook as JSON with baked-in outputs, which invites out-of-order execution bugs, stale variables, and unreadable diffs. Marimo derives execution order from a dataflow DAG, forbids defining the same variable in two cells, and saves the notebook as a plain Python file. The practical effect is that a Marimo notebook is reproducible by construction, whereas a Jupyter notebook is reproducible only if you remember to "Restart &amp;amp; Run All."&lt;/p&gt;

&lt;h3&gt;
  
  
  How does Marimo know which cells to rerun?
&lt;/h3&gt;

&lt;p&gt;Marimo statically analyzes each cell to see which global variables it defines and which it references, then draws an edge from the defining cell to the referencing cell. When a cell changes, Marimo reruns that cell and its transitive descendants in the graph — nothing upstream and nothing on an unrelated branch. Deleting a cell removes its variables and invalidates the cells that used them, so no cell keeps running against a ghost value.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why are Marimo notebooks stored as &lt;code&gt;.py&lt;/code&gt; files?
&lt;/h3&gt;

&lt;p&gt;Because a plain Python file is diffable, reviewable, importable, and executable without a kernel. A &lt;code&gt;.ipynb&lt;/code&gt; is JSON that mixes source, outputs, and execution counts, producing noisy diffs and merge conflicts and letting stale outputs live on disk. A Marimo &lt;code&gt;.py&lt;/code&gt; file contains only code, so one edit is one line of diff, and the same file runs as an interactive app (&lt;code&gt;marimo run&lt;/code&gt;) or a headless script (&lt;code&gt;python notebook.py&lt;/code&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  How do Marimo UI elements work without callbacks?
&lt;/h3&gt;

&lt;p&gt;A UI element such as &lt;code&gt;mo.ui.slider&lt;/code&gt; is a Python object with a live &lt;code&gt;.value&lt;/code&gt;. You define and display it in one cell, and any &lt;em&gt;other&lt;/em&gt; cell that reads its &lt;code&gt;.value&lt;/code&gt; automatically becomes a dependent in the DAG. Interacting with the widget is treated as a change to its cell, so Marimo reruns only the dependent cells — there is no &lt;code&gt;on_change&lt;/code&gt; handler and no global mutation, because the dependency edge already declares what to recompute.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can Marimo run SQL?
&lt;/h3&gt;

&lt;p&gt;Yes — Marimo has first-class SQL cells backed by an embedded &lt;strong&gt;DuckDB&lt;/strong&gt; engine, invoked through &lt;code&gt;mo.sql(...)&lt;/code&gt;. A SQL cell can query your in-memory Python dataframes by variable name (no load step) and returns the result as a dataframe that is itself a reactive variable. You can interpolate Python values, including a UI element's &lt;code&gt;.value&lt;/code&gt;, straight into the query with an f-string, so a slider can drive a live SQL result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every Marimo idea above, from the reactive dataflow DAG and the single-definition rule to UI elements bound by `.value` and SQL-over-dataframe with DuckDB, maps to a hands-on practice room where you build the analysis against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you make this notebook reproducible?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;Practice data-analysis problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Pipeline-design drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Papermill &amp; Notebook Pipelines: Parametrized, Scheduled, Version-Controlled Notebooks</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Thu, 24 Sep 2026 17:43:33 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/papermill-notebook-pipelines-parametrized-scheduled-version-controlled-notebooks-4gb9</link>
      <guid>https://dev.to/gowthampotureddi/papermill-notebook-pipelines-parametrized-scheduled-version-controlled-notebooks-4gb9</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;papermill notebook pipelines&lt;/code&gt;&lt;/strong&gt; are what you get when you stop treating a Jupyter notebook as a throwaway scratchpad and start treating it as a runnable, parametrized job. Papermill is a small open-source tool — a Python library and a command-line program — that takes a notebook, injects a fresh set of parameters into it, runs every cell top to bottom, and writes a &lt;em&gt;new&lt;/em&gt; notebook that contains the code, the injected values, and every output the run produced. The template is never mutated; each execution leaves behind its own fully-rendered artifact.&lt;/p&gt;

&lt;p&gt;That one move changes the economics of the work an analyst already does in a notebook. Instead of copying a notebook, hand-editing the date at the top, running it, and screenshotting a chart into an email, you run &lt;code&gt;papermill report.ipynb out/2026-09-15.ipynb -p run_date 2026-09-15&lt;/code&gt; and the executed notebook — charts, tables, logs, and all — becomes the deliverable. This guide walks the five things an interviewer will actually probe: the &lt;code&gt;parameters&lt;/code&gt; cell tag and &lt;code&gt;papermill.execute_notebook&lt;/code&gt;, the executed output notebook as an artifact plus &lt;code&gt;nbconvert&lt;/code&gt;, collecting structured results with &lt;code&gt;scrapbook.glue&lt;/code&gt;, orchestrating notebooks on a schedule with Airflow's &lt;code&gt;PapermillOperator&lt;/code&gt;, and version-controlling notebooks with &lt;code&gt;nbdime&lt;/code&gt;. Each section pairs the concept with a Solution-Tail interview answer: code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works — including the honest part, which is when notebooks-as-pipelines is the wrong tool.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgf3obbrwx54r6w4wmup1.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgf3obbrwx54r6w4wmup1.jpeg" alt="PipeCode blog header for Papermill notebook pipelines — bold white headline 'Papermill: Notebook Pipelines' with subtitle 'Parametrized · Scheduled · Version-Controlled' and a stylised parametrized-notebook-to-executed-artifact scene on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;pipeline-design practice library →&lt;/a&gt;, rehearse the run-cadence decisions on the &lt;a href="https://pipecode.ai/explore/practice/topic/scheduling" rel="noopener noreferrer"&gt;scheduling practice set →&lt;/a&gt;, and shape the parametrized outputs on the &lt;a href="https://pipecode.ai/explore/practice/topic/data-transformation" rel="noopener noreferrer"&gt;data-transformation practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why notebooks-as-pipelines — Papermill's job and its limits&lt;/li&gt;
&lt;li&gt;The parameters cell tag &amp;amp; execute_notebook&lt;/li&gt;
&lt;li&gt;The executed output notebook as an artifact&lt;/li&gt;
&lt;li&gt;Collecting outputs with scrapbook &amp;amp; glue&lt;/li&gt;
&lt;li&gt;Orchestrating &amp;amp; version-controlling notebook pipelines&lt;/li&gt;
&lt;li&gt;Cheat sheet — Papermill recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why notebooks-as-pipelines — Papermill's job and its limits
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Papermill parametrizes and executes a notebook, producing a second notebook — that one fact decides where it fits
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;Papermill runs a notebook the way a function runs a body — you pass parameters in, it executes, and it returns a new notebook with the results baked in&lt;/strong&gt;. The input notebook is read-only during a run; the value Papermill produces is an &lt;em&gt;output notebook&lt;/em&gt;, an ordinary &lt;code&gt;.ipynb&lt;/code&gt; in which every cell has been executed and its outputs captured. Everything else — scheduling, storing to S3, rendering to HTML — is built on top of that single behaviour.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What Papermill actually is (and is not).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A parametrizer.&lt;/strong&gt; Papermill's headline feature is injecting parameters into a notebook without editing the source. You mark one cell, and Papermill overrides its variables at run time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An executor.&lt;/strong&gt; It runs the notebook cell-by-cell against a Jupyter kernel (&lt;code&gt;python3&lt;/code&gt; by default, but any installed kernel), saving outputs and cell timings as it goes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not a scheduler.&lt;/strong&gt; Papermill does not have a cron, a DAG, or a UI. You invoke it from a script, a CI job, or an orchestrator. "Scheduled notebooks" means &lt;em&gt;something else&lt;/em&gt; triggers &lt;code&gt;papermill&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not a transformation framework.&lt;/strong&gt; Papermill neither knows nor cares what the notebook does; there is no dependency graph between cells beyond top-to-bottom order.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The good-idea case — when the artifact is the point.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The output notebook is a self-documenting audit log.&lt;/strong&gt; Inputs, code, logs, tables, and charts live in one file for one run — reproducible and reviewable months later without re-running anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The author is the operator.&lt;/strong&gt; An analyst who already lives in Jupyter can promote their notebook to a scheduled job with zero rewrite into a "real" script, which is often the difference between a report shipping and not shipping.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parametrized fan-out is trivial.&lt;/strong&gt; Run the same notebook once per region, per date, or per customer by looping over parameter sets — each run is an independent artifact.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The anti-pattern case — when to reach for a script or task instead.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hidden execution state.&lt;/strong&gt; Notebooks invite out-of-order cell runs and lingering variables; a notebook that only works if you run cell 7 before cell 3 is a landmine. Papermill always runs top-to-bottom, which helps, but complex control flow still belongs in tested modules.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heavy DAGs and shared logic.&lt;/strong&gt; If tasks fan out, retry independently, or share a library, a notebook is the wrong container — extract the logic into a package and call it from Airflow tasks or a script.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team-edited notebooks.&lt;/strong&gt; Concurrent edits to a JSON notebook merge badly; long-lived, multi-author pipeline logic wants plain &lt;code&gt;.py&lt;/code&gt; files under review.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you say &lt;strong&gt;"Papermill produces a new executed notebook, it does not mutate the input"&lt;/strong&gt; early? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you separate &lt;strong&gt;"Papermill executes" from "something else schedules"&lt;/strong&gt;? — required framing.&lt;/li&gt;
&lt;li&gt;Can you name &lt;strong&gt;both the good case (the artifact is an audit log) and the anti-pattern (hidden state, heavy DAGs)&lt;/strong&gt; unprompted? — the maturity signal.&lt;/li&gt;
&lt;li&gt;Do you mention that &lt;strong&gt;parameters are injected via a tagged cell&lt;/strong&gt;, not by editing source? — the whole mechanism.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — five lines that turn a notebook into a job
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The canonical Papermill "hello world" takes an existing notebook and runs it with one overridden parameter, writing the executed result somewhere new. It looks trivial, and that is the point: the same call that runs a toy report scales unchanged to a nightly job, because Papermill only ever sees "a notebook, some parameters, an output path." No cell is edited by hand; the override is injected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Run a &lt;code&gt;report.ipynb&lt;/code&gt; for a specific date and land the executed notebook at a new path, without touching the original.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;argument&lt;/th&gt;
&lt;th&gt;value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;input notebook&lt;/td&gt;
&lt;td&gt;&lt;code&gt;report.ipynb&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;output notebook&lt;/td&gt;
&lt;td&gt;&lt;code&gt;runs/report-2026-09-15.ipynb&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;parameter &lt;code&gt;run_date&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;2026-09-15&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;papermill&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pm&lt;/span&gt;

&lt;span class="n"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute_notebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;report.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                       &lt;span class="c1"&gt;# input template (never modified)
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;runs/report-2026-09-15.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="c1"&gt;# executed output artifact
&lt;/span&gt;    &lt;span class="n"&gt;parameters&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2026-09-15&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;execute_notebook&lt;/code&gt; reads &lt;code&gt;report.ipynb&lt;/code&gt;, finds the cell tagged &lt;code&gt;parameters&lt;/code&gt;, and inserts a new &lt;code&gt;injected-parameters&lt;/code&gt; cell directly beneath it that sets &lt;code&gt;run_date = "2026-09-15"&lt;/code&gt;, overriding the default. It then launches the &lt;code&gt;python3&lt;/code&gt; kernel and runs every cell in order, capturing each cell's outputs. Finally it writes the fully-executed notebook to &lt;code&gt;runs/report-2026-09-15.ipynb&lt;/code&gt;. The original &lt;code&gt;report.ipynb&lt;/code&gt; is untouched — you could run it a hundred times with a hundred dates and get a hundred artifacts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Papermill produced&lt;/th&gt;
&lt;th&gt;value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;input notebook&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;report.ipynb&lt;/code&gt; (unchanged)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;output notebook&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;runs/report-2026-09-15.ipynb&lt;/code&gt; (all cells executed)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;injected cell&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;injected-parameters&lt;/code&gt; with &lt;code&gt;run_date = "2026-09-15"&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;captured&lt;/td&gt;
&lt;td&gt;every cell's outputs + per-cell execution time&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If the work already lives in a notebook and the deliverable is "this notebook, run with today's inputs," Papermill is the smallest possible upgrade from manual to reproducible — one &lt;code&gt;execute_notebook&lt;/code&gt; call, no rewrite.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. The parameters cell tag &amp;amp; execute_notebook
&lt;/h2&gt;

&lt;h3&gt;
  
  
  One tagged cell is the whole parameter interface — Papermill injects overrides right beneath it
&lt;/h3&gt;

&lt;p&gt;Papermill has exactly one mechanism for parametrizing a notebook, and an interviewer who asks "how does Papermill inject parameters?" wants it precisely. You tag &lt;strong&gt;one&lt;/strong&gt; cell with the cell tag &lt;code&gt;parameters&lt;/code&gt;. That cell holds default assignments. At run time Papermill inserts a second, machine-generated cell — labelled &lt;code&gt;injected-parameters&lt;/code&gt; — immediately after it, containing the values you passed. Because Python executes top-to-bottom, the injected assignments run after the defaults and win.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tagging mechanism.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The &lt;code&gt;parameters&lt;/code&gt; tag.&lt;/strong&gt; In JupyterLab you add it via &lt;em&gt;Property Inspector → Cell Tags → add tag &lt;code&gt;parameters&lt;/code&gt;&lt;/em&gt;. It is metadata on the cell, not a comment; Papermill looks for exactly this tag.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Defaults belong in the tagged cell.&lt;/strong&gt; Assign every parameter a sensible default there (&lt;code&gt;run_date = "2026-01-01"&lt;/code&gt;, &lt;code&gt;region = "us"&lt;/code&gt;). This keeps the notebook runnable standalone in Jupyter with no Papermill involved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The injected cell.&lt;/strong&gt; Papermill writes &lt;code&gt;injected-parameters&lt;/code&gt; right below the tagged cell. If no tagged cell exists, Papermill injects at the top and warns — your defaults then never run, a common footgun.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Passing parameters — library and CLI.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Library.&lt;/strong&gt; &lt;code&gt;pm.execute_notebook(input, output, parameters={"region": "eu", "limit": 500})&lt;/code&gt; passes a dict; types are preserved (int stays int, list stays list).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CLI &lt;code&gt;-p&lt;/code&gt; (typed).&lt;/strong&gt; &lt;code&gt;papermill in.ipynb out.ipynb -p limit 500 -p region eu&lt;/code&gt; — Papermill infers &lt;code&gt;500&lt;/code&gt; as an int and &lt;code&gt;eu&lt;/code&gt; as a string.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CLI &lt;code&gt;-r&lt;/code&gt; (raw string).&lt;/strong&gt; &lt;code&gt;-r zip 02139&lt;/code&gt; forces a string, so a zero-padded code is not mangled into the integer &lt;code&gt;2139&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CLI &lt;code&gt;-y&lt;/code&gt; / &lt;code&gt;-f&lt;/code&gt; (YAML / file).&lt;/strong&gt; &lt;code&gt;-y "region: eu\nskus: [a, b]"&lt;/code&gt; passes inline YAML; &lt;code&gt;-f params.yaml&lt;/code&gt; reads a file — the clean way to pass nested or list parameters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;-k&lt;/code&gt; kernel.&lt;/strong&gt; &lt;code&gt;-k python3&lt;/code&gt; (or an R / Scala kernel) picks the kernel to execute against.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why injection beats editing source.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The template file is never modified, so it stays clean in version control and safe to run concurrently with different parameter sets.&lt;/li&gt;
&lt;li&gt;Types survive: a dict passed in Python arrives as a dict, not a re-parsed string.&lt;/li&gt;
&lt;li&gt;The injected cell is visible in the output notebook, so anyone reading the artifact sees exactly which values produced it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0xm16j3j89ufe7ac7kx0.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0xm16j3j89ufe7ac7kx0.jpeg" alt="Iconographic Papermill parametrize diagram — a template notebook with a cell tagged parameters holding defaults, Papermill injecting an injected-parameters cell just below it with overridden values, and execute_notebook producing an output notebook." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — a parameters cell with defaults, overridden at run time
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday pattern is a first code cell tagged &lt;code&gt;parameters&lt;/code&gt; that declares every input with a default, followed by cells that use those names. Papermill overrides a subset; the rest keep their defaults. This is what lets the same notebook run interactively (defaults) and as a job (overrides).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A notebook's first cell is tagged &lt;code&gt;parameters&lt;/code&gt; with &lt;code&gt;region = "us"&lt;/code&gt; and &lt;code&gt;limit = 100&lt;/code&gt;. You run it with Papermill passing &lt;code&gt;region="eu"&lt;/code&gt; only. What values do the downstream cells see?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;region&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;    &lt;span class="c1"&gt;# first cell of report.ipynb, tagged "parameters"
&lt;/span&gt;&lt;span class="n"&gt;limit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;papermill&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pm&lt;/span&gt;

&lt;span class="n"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute_notebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;report.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;runs/report-eu.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;parameters&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;        &lt;span class="c1"&gt;# override region; limit keeps its default
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The CLI form is identical: &lt;code&gt;papermill report.ipynb runs/report-eu.ipynb -p region eu&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Papermill locates the &lt;code&gt;parameters&lt;/code&gt;-tagged cell and inserts &lt;code&gt;injected-parameters&lt;/code&gt; right after it containing &lt;code&gt;region = "eu"&lt;/code&gt;. When the notebook runs, the tagged cell sets &lt;code&gt;region = "us"&lt;/code&gt; and &lt;code&gt;limit = 100&lt;/code&gt;, then the injected cell immediately reassigns &lt;code&gt;region = "eu"&lt;/code&gt;. Because &lt;code&gt;limit&lt;/code&gt; was not passed, it keeps its default &lt;code&gt;100&lt;/code&gt;. Every downstream cell reads &lt;code&gt;region == "eu"&lt;/code&gt; and &lt;code&gt;limit == 100&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;variable&lt;/th&gt;
&lt;th&gt;source of final value&lt;/th&gt;
&lt;th&gt;value at run time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;region&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;injected-parameters (override)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;"eu"&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;limit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;parameters cell (default)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;100&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Give every parameter a default in the tagged cell, then override only what changes per run — the notebook stays runnable by hand and precise as a job.&lt;/p&gt;

&lt;h3&gt;
  
  
  Papermill interview question on parameter injection
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A teammate reports that their Papermill run "ignores the parameters" — the notebook always uses the defaults no matter what &lt;code&gt;-p&lt;/code&gt; values they pass. The notebook opens fine and runs top-to-bottom in Jupyter. What is almost certainly wrong, and how would you prove and fix it?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using the parameters cell tag correctly
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;papermill&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pm&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;nbformat&lt;/span&gt;

&lt;span class="n"&gt;nb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;nbformat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;report.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;as_version&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;tagged&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;                                       &lt;span class="c1"&gt;# prove it: which cells carry the tag?
&lt;/span&gt;    &lt;span class="n"&gt;idx&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;idx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cells&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;parameters&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;metadata&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tags&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cells tagged &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;parameters&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tagged&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="c1"&gt;# [] means none — that's the bug
&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;tagged&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;                                   &lt;span class="c1"&gt;# fix: tag the defaults cell, then re-run
&lt;/span&gt;    &lt;span class="n"&gt;nb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cells&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;metadata&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;setdefault&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tags&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]).&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;parameters&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;nbformat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nb&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;report.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute_notebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;report.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;runs/report-eu.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;parameters&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;condition&lt;/th&gt;
&lt;th&gt;Papermill behaviour&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;no cell tagged &lt;code&gt;parameters&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;injects &lt;code&gt;injected-parameters&lt;/code&gt; at top, emits a warning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;injected cell runs first&lt;/td&gt;
&lt;td&gt;later default cell reassigns &lt;code&gt;region = "us"&lt;/code&gt;, clobbering the override&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;tag the defaults cell&lt;/td&gt;
&lt;td&gt;injected cell now lands &lt;em&gt;after&lt;/em&gt; defaults&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;re-run with &lt;code&gt;-p region eu&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;override runs last and wins → &lt;code&gt;region == "eu"&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Papermill only overrides variables via the &lt;code&gt;injected-parameters&lt;/code&gt; cell it inserts &lt;strong&gt;right after&lt;/strong&gt; the &lt;code&gt;parameters&lt;/code&gt;-tagged cell.&lt;/li&gt;
&lt;li&gt;With no tagged cell, Papermill still injects — but at the &lt;strong&gt;top&lt;/strong&gt; — so any later plain cell that reassigns &lt;code&gt;region&lt;/code&gt; overwrites the injected value; the defaults appear to "win."&lt;/li&gt;
&lt;li&gt;The warning &lt;code&gt;No cell tagged 'parameters'&lt;/code&gt; in the output notebook is the diagnostic; the empty &lt;code&gt;tagged&lt;/code&gt; list is the proof.&lt;/li&gt;
&lt;li&gt;Tagging the defaults cell places the injected cell after the defaults, so the override is the last assignment and takes effect.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;state&lt;/th&gt;
&lt;th&gt;tagged cells&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;region&lt;/code&gt; at run time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;before fix&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;"us"&lt;/code&gt; (default wins)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;after fix&lt;/td&gt;
&lt;td&gt;cell 0&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;"eu"&lt;/code&gt; (override wins)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cell tag, not comment&lt;/strong&gt;&lt;/strong&gt; — Papermill keys off notebook metadata &lt;code&gt;tags: ["parameters"]&lt;/code&gt;; a &lt;code&gt;# parameters&lt;/code&gt; comment does nothing, which is the single most common Papermill mistake.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Injection position&lt;/strong&gt;&lt;/strong&gt; — overrides are inserted &lt;em&gt;after&lt;/em&gt; the tagged cell so they run last; with no tag, injection at the top is beaten by later reassignments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Defaults keep it runnable&lt;/strong&gt;&lt;/strong&gt; — the tagged cell's defaults let the notebook run in plain Jupyter, while Papermill supplies production values.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Inspect to diagnose&lt;/strong&gt;&lt;/strong&gt; — reading &lt;code&gt;cell.metadata.tags&lt;/code&gt; with &lt;code&gt;nbformat&lt;/code&gt; turns a vague "it ignores params" into a one-line proof.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — injection and tag lookup are O(cells), trivial against the runtime of the notebook's actual work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Pipelines&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — pipelines&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Parametrized pipeline-design problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Transform&lt;/span&gt;
&lt;span&gt;Topic — data-transformation&lt;/span&gt;
&lt;strong&gt;Parameter-driven transformation problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-transformation" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. The executed output notebook as an artifact
&lt;/h2&gt;
&lt;h3&gt;
  
  
  The output notebook is the deliverable — a fully-run, self-contained record of one execution
&lt;/h3&gt;

&lt;p&gt;The feature that makes Papermill a &lt;em&gt;pipeline&lt;/em&gt; tool rather than a runner is that &lt;strong&gt;every execution produces a durable, human-readable artifact: the output notebook&lt;/strong&gt;. The input template stays pristine; the output is the same notebook with every cell run, every output captured, per-cell timing recorded, and the injected parameters visible at the top. Months later you can open it and see exactly what code ran, on what inputs, with what result — no re-execution, no guessing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What lands in the output notebook.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Executed cells with outputs.&lt;/strong&gt; Tables, stdout, logs, and rendered charts (as embedded images) are all stored inline, so the file is portable and complete.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Execution metadata.&lt;/strong&gt; Papermill records per-cell start/end times and a run status in the notebook metadata, so you can see which cell was slow or where it failed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The injected-parameters cell.&lt;/strong&gt; The exact values that produced this run are right there in the artifact — the run is self-describing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Errors are captured, not hidden.&lt;/strong&gt; By default a cell that raises stops the run and the traceback is written into the output notebook; the file is saved so you can debug the failure post-mortem.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Where the output can live.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Path-based backends.&lt;/strong&gt; The output path can be local, &lt;code&gt;s3://bucket/key.ipynb&lt;/code&gt;, &lt;code&gt;gs://bucket/key.ipynb&lt;/code&gt;, or Azure Blob — Papermill uses &lt;code&gt;fsspec&lt;/code&gt;-style handlers, so the same call writes to object storage with no code change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-run paths.&lt;/strong&gt; Templating the date or run id into the output path (&lt;code&gt;runs/report-{date}.ipynb&lt;/code&gt;) keeps every execution's artifact instead of overwriting one file.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Turning the artifact into something to share.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;--report-mode&lt;/code&gt;.&lt;/strong&gt; Runs Papermill with input cells hidden in the rendered result, so stakeholders see outputs (charts, tables) without the code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;jupyter nbconvert&lt;/code&gt;.&lt;/strong&gt; Converts the executed notebook to HTML or PDF: &lt;code&gt;jupyter nbconvert --to html runs/report.ipynb&lt;/code&gt;. This is the standard "email a clean report" step — nbconvert renders, Papermill executes; they compose.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;--no-input&lt;/code&gt;.&lt;/strong&gt; An nbconvert flag that strips code cells from the rendered HTML/PDF for a purely visual report.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe197ny929munvv9do3oz.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe197ny929munvv9do3oz.jpeg" alt="Iconographic Papermill output-artifact diagram — the untouched input notebook, a fully executed output notebook with cell outputs, timings and status, storage backends for the output path, and nbconvert rendering it to HTML and PDF." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — execute to S3, then render to HTML
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; A production pattern is: execute the notebook and land the artifact in object storage keyed by run date, then render a human-friendly HTML from that executed notebook. The execute step preserves the machine-readable record; the nbconvert step produces the shareable document. Neither step touches the template.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Run &lt;code&gt;report.ipynb&lt;/code&gt; for &lt;code&gt;2026-09-15&lt;/code&gt;, store the executed notebook to S3 under a dated key, and produce a code-free HTML report from it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;tool&lt;/th&gt;
&lt;th&gt;target&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;execute&lt;/td&gt;
&lt;td&gt;papermill&lt;/td&gt;
&lt;td&gt;&lt;code&gt;s3://reports/2026-09-15/report.ipynb&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;render&lt;/td&gt;
&lt;td&gt;nbconvert&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;report-2026-09-15.html&lt;/code&gt; (no code)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;papermill&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pm&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;

&lt;span class="n"&gt;out_nb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s3://reports/2026-09-15/report.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute_notebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;report.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;out_nb&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;parameters&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2026-09-15&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;                            &lt;span class="c1"&gt;# render the executed notebook to code-free HTML
&lt;/span&gt;    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jupyter&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nbconvert&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--to&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;html&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--no-input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
     &lt;span class="n"&gt;out_nb&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;report-2026-09-15&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;execute_notebook&lt;/code&gt; runs the parametrized notebook and writes the executed artifact straight to S3 — the &lt;code&gt;s3://&lt;/code&gt; prefix routes the write through the object-store handler, no extra upload code. The executed notebook contains every output. &lt;code&gt;jupyter nbconvert --to html --no-input&lt;/code&gt; then reads that executed notebook and emits &lt;code&gt;report-2026-09-15.html&lt;/code&gt; containing only the rendered outputs (charts, tables), because &lt;code&gt;--no-input&lt;/code&gt; drops the code cells. The template &lt;code&gt;report.ipynb&lt;/code&gt; remains unchanged.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;artifact&lt;/th&gt;
&lt;th&gt;contents&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;s3://reports/2026-09-15/report.ipynb&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;executed notebook: code + outputs + injected params + timings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;report-2026-09-15.html&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;rendered outputs only (no code), ready to email&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Execute once to a dated, durable path for the audit trail; convert that same artifact with nbconvert for humans — never re-run the notebook just to get a different format.&lt;/p&gt;

&lt;h3&gt;
  
  
  Papermill interview question on failure artifacts
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Your nightly Papermill job failed at cell 12 of 20. A teammate says "the run is gone, we have to reproduce it locally." Are they right? What did Papermill leave behind, and how do you use it?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using the saved output notebook on failure
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;papermill&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pm&lt;/span&gt;

&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute_notebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;etl.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;runs/etl-2026-09-15.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;     &lt;span class="c1"&gt;# saved even if a cell raises
&lt;/span&gt;        &lt;span class="n"&gt;parameters&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2026-09-15&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;exceptions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;PapermillExecutionError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# The output notebook is already on disk with the traceback in cell 12.
&lt;/span&gt;    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;failed cell:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cell_index&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ename&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt;                                 &lt;span class="c1"&gt;# let the orchestrator mark the task failed
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;event&lt;/th&gt;
&lt;th&gt;what Papermill does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;cells 1–11 run&lt;/td&gt;
&lt;td&gt;outputs captured into the output notebook in progress&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;cell 12 raises&lt;/td&gt;
&lt;td&gt;traceback written into cell 12's output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;run halts&lt;/td&gt;
&lt;td&gt;output notebook &lt;strong&gt;saved&lt;/strong&gt; to &lt;code&gt;runs/etl-2026-09-15.ipynb&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;&lt;code&gt;PapermillExecutionError&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;raised to the caller with &lt;code&gt;cell_index=12&lt;/code&gt; and error name&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Papermill executes cells in order, persisting outputs as it goes, so partial progress is not lost.&lt;/li&gt;
&lt;li&gt;When cell 12 raises, Papermill records the full traceback &lt;em&gt;inside&lt;/em&gt; the output notebook rather than discarding it.&lt;/li&gt;
&lt;li&gt;It then saves the output notebook to the specified path — the failure artifact exists on disk even though the run failed.&lt;/li&gt;
&lt;li&gt;It re-raises &lt;code&gt;PapermillExecutionError&lt;/code&gt; so the orchestrator sees a failed task, while you open the saved notebook to debug the exact failing cell with its inputs intact.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;claim&lt;/th&gt;
&lt;th&gt;reality&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"the run is gone"&lt;/td&gt;
&lt;td&gt;false — output notebook saved with the traceback&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;where to debug&lt;/td&gt;
&lt;td&gt;open &lt;code&gt;runs/etl-2026-09-15.ipynb&lt;/code&gt;, jump to cell 12&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Fail-with-artifact&lt;/strong&gt;&lt;/strong&gt; — Papermill saves the output notebook even on error, so a failed run is a debuggable record, not a void.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Traceback in place&lt;/strong&gt;&lt;/strong&gt; — the exception is captured in the failing cell's output, so you see the error next to the code and the injected parameters that triggered it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;PapermillExecutionError&lt;/strong&gt;&lt;/strong&gt; — a typed exception carrying &lt;code&gt;cell_index&lt;/code&gt; and error name lets orchestration fail cleanly and alert precisely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Reproducibility for free&lt;/strong&gt;&lt;/strong&gt; — the injected-parameters cell means re-running the exact scenario is copy-paste, no guessing which date failed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — persisting outputs incrementally is O(output size); the debugging time it saves is the real payoff.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Pipelines&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — pipelines&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Artifact-and-observability pipeline problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Transform&lt;/span&gt;
&lt;span&gt;Topic — data-transformation&lt;/span&gt;
&lt;strong&gt;Output-shaping and reporting problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-transformation" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Collecting outputs with scrapbook &amp;amp; glue
&lt;/h2&gt;
&lt;h3&gt;
  
  
  &lt;code&gt;scrapbook.glue&lt;/code&gt; writes named values into the notebook, &lt;code&gt;read_notebook&lt;/code&gt; reads them back — data escapes the notebook without a side database
&lt;/h3&gt;

&lt;p&gt;An executed notebook is great for humans, but a pipeline usually needs to pull &lt;em&gt;structured&lt;/em&gt; results back out — a row count, a model score, a small DataFrame — so a downstream step can use it. Papermill's original &lt;code&gt;record&lt;/code&gt; / &lt;code&gt;read_notebook&lt;/code&gt; API is deprecated; the supported way is the companion library &lt;strong&gt;scrapbook&lt;/strong&gt;. The mechanism is symmetric: inside the notebook you &lt;code&gt;sb.glue("name", value)&lt;/code&gt; to persist a value, and outside you &lt;code&gt;sb.read_notebook(path).scraps&lt;/code&gt; to read it back. The values ("scraps") are stored inside the output notebook itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The glue side (inside the notebook).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;sb.glue("rows", 1240)&lt;/code&gt;.&lt;/strong&gt; Persists a JSON-serializable value under a name; it is stored in the cell's metadata as a scrap and survives in the output notebook.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;sb.glue("summary", df, encoder="pandas")&lt;/code&gt;.&lt;/strong&gt; Encoders let you glue richer objects (a DataFrame, an Arrow table) that round-trip back to the same type.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;sb.glue("chart", fig, display=True)&lt;/code&gt;.&lt;/strong&gt; With &lt;code&gt;display=True&lt;/code&gt; the value is &lt;em&gt;also&lt;/em&gt; rendered visibly in the notebook, so it is both machine-readable data and a human-visible output.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The read side (downstream code).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;sb.read_notebook(path).scraps&lt;/code&gt;.&lt;/strong&gt; Returns a name → scrap mapping; &lt;code&gt;.scraps["rows"].data&lt;/code&gt; is the original value with its type restored.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;sb.read_notebooks(dir).scraps_report()&lt;/code&gt;&lt;/strong&gt; / &lt;strong&gt;&lt;code&gt;.papermill_dataframe&lt;/code&gt;.&lt;/strong&gt; Read a whole directory of executed notebooks and assemble one row per run — exactly what you want after a parametrized fan-out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Type fidelity.&lt;/strong&gt; Because scraps carry an encoder tag, a glued DataFrame comes back as a DataFrame, not a re-parsed string.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why this beats the alternatives.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No side channel.&lt;/strong&gt; The result travels &lt;em&gt;with&lt;/em&gt; the artifact — you do not need a separate table, file, or return-value plumbing to move a scalar from a notebook to the next step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batch-friendly.&lt;/strong&gt; After running one notebook per region, a single &lt;code&gt;read_notebooks&lt;/code&gt; call collects every region's metrics into a DataFrame for comparison.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auditable.&lt;/strong&gt; The glued value is visible in the same artifact that shows the code and inputs that produced it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffl2twnwazvpkdvdtrlfj.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffl2twnwazvpkdvdtrlfj.jpeg" alt="Iconographic scrapbook diagram — sb.glue persisting named values into the executed notebook as scraps, and sb.read_notebook reading the scraps back out into a downstream DataFrame across a batch of runs." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — glue a metric, read it in the next step
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday scrapbook pattern is: the notebook computes something and glues the number it wants to expose; an orchestrating script runs the notebook, then reads that number back to decide what to do next. This turns a notebook into a callable step with a return value.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A notebook computes &lt;code&gt;row_count&lt;/code&gt;. Glue it, run the notebook with Papermill, then read &lt;code&gt;row_count&lt;/code&gt; back in the orchestrator and branch on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;scrapbook&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;sb&lt;/span&gt;        &lt;span class="c1"&gt;# inside etl.ipynb, after the load cell
&lt;/span&gt;&lt;span class="n"&gt;row_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;          &lt;span class="c1"&gt;# e.g. 1240
&lt;/span&gt;&lt;span class="n"&gt;sb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;glue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;row_count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;row_count&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;papermill&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pm&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;scrapbook&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;sb&lt;/span&gt;

&lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;runs/etl-2026-09-15.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute_notebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;etl.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;parameters&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2026-09-15&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="n"&gt;nb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_notebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                   &lt;span class="c1"&gt;# read the glued value back out
&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;nb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;scraps&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;row_count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;           &lt;span class="c1"&gt;# -&amp;gt; 1240, as an int
&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;load produced zero rows — failing the pipeline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;loaded &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; rows&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;actor&lt;/th&gt;
&lt;th&gt;effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;notebook&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;sb.glue("row_count", 1240)&lt;/code&gt; stores a scrap in the output notebook&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;papermill&lt;/td&gt;
&lt;td&gt;writes the executed notebook (with the scrap) to &lt;code&gt;out&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;orchestrator&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;sb.read_notebook(out).scraps["row_count"].data&lt;/code&gt; → &lt;code&gt;1240&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;orchestrator&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;rows != 0&lt;/code&gt;, so the pipeline proceeds&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Inside the notebook, &lt;code&gt;glue&lt;/code&gt; serializes &lt;code&gt;row_count&lt;/code&gt; and stores it as a named scrap in that cell's metadata.&lt;/li&gt;
&lt;li&gt;Papermill saves the executed notebook, so the scrap travels inside the artifact — no external write.&lt;/li&gt;
&lt;li&gt;The orchestrator reads the notebook back and pulls &lt;code&gt;scraps["row_count"].data&lt;/code&gt;, recovering the integer &lt;code&gt;1240&lt;/code&gt; with its type intact.&lt;/li&gt;
&lt;li&gt;Because the value is a real int, the &lt;code&gt;== 0&lt;/code&gt; guard works directly, and the pipeline branches on genuine notebook output.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;scrap&lt;/th&gt;
&lt;th&gt;stored in&lt;/th&gt;
&lt;th&gt;recovered value&lt;/th&gt;
&lt;th&gt;type&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;row_count&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;runs/etl-2026-09-15.ipynb&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1240&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;int&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;glue = persist-with-artifact&lt;/strong&gt;&lt;/strong&gt; — the value is written into the output notebook, so results and evidence never drift apart.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;scraps carry types&lt;/strong&gt;&lt;/strong&gt; — an encoder tag round-trips ints, dicts, and DataFrames back to their original types instead of strings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;read_notebook as return value&lt;/strong&gt;&lt;/strong&gt; — the downstream step treats the notebook like a function that returned data, enabling real control flow (guards, branching).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;No side database&lt;/strong&gt;&lt;/strong&gt; — moving a scalar between steps needs no extra table or file, cutting a whole class of plumbing bugs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — glue/read are O(value size); for scalars and small frames it is negligible next to the notebook's compute.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Transform&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — data-transformation&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Extract-and-collect result problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-transformation" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Pipelines&lt;/span&gt;
&lt;span&gt;Topic — pipelines&lt;/span&gt;
&lt;strong&gt;Step-to-step handoff pipeline problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Orchestrating &amp;amp; version-controlling notebook pipelines
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Something else schedules Papermill, and nbdime tames the diff — the two things that make notebooks production-grade
&lt;/h3&gt;

&lt;p&gt;Papermill executes; it does not schedule and it does not version-control. Making notebook pipelines production-grade means solving those two separately: an orchestrator (usually Airflow) triggers &lt;code&gt;papermill&lt;/code&gt; on a cadence with per-run parameters, and a notebook-aware diff tool (&lt;code&gt;nbdime&lt;/code&gt;) plus output stripping (&lt;code&gt;nbstripout&lt;/code&gt;) make the JSON reviewable in git. Say it in one breath: &lt;strong&gt;Airflow's &lt;code&gt;PapermillOperator&lt;/code&gt; runs the notebook on a schedule, and nbdime/nbstripout keep the source diffable&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scheduling with Airflow.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;PapermillOperator&lt;/code&gt;.&lt;/strong&gt; From &lt;code&gt;airflow.providers.papermill&lt;/code&gt;, it wraps &lt;code&gt;execute_notebook&lt;/code&gt;: you give it &lt;code&gt;input_nb&lt;/code&gt;, &lt;code&gt;output_nb&lt;/code&gt;, and &lt;code&gt;parameters&lt;/code&gt;, and it becomes a task in a DAG.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Templated per-run paths and dates.&lt;/strong&gt; Use Jinja templating (&lt;code&gt;{{ ds }}&lt;/code&gt;) so each scheduled run writes its own dated output notebook and passes the run date as a parameter — one artifact per run, no overwrites.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Idempotent parameters.&lt;/strong&gt; Parametrize by the &lt;em&gt;logical&lt;/em&gt; run date, not &lt;code&gt;datetime.now()&lt;/code&gt;, so a re-run of the same interval reproduces the same result — the property that makes retries and backfills safe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operator-level retries.&lt;/strong&gt; Retries, alerts, and SLAs live on the Airflow task, not in the notebook; Papermill just runs the notebook and fails loudly if a cell raises.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Version-controlling notebooks.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The problem.&lt;/strong&gt; A &lt;code&gt;.ipynb&lt;/code&gt; is JSON containing source, outputs, execution counts, and metadata, so a one-line code change shows up as a huge, unreadable git diff full of base64 image blobs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;nbstripout&lt;/code&gt;.&lt;/strong&gt; A git filter that strips outputs and execution counts on commit, so the repo stores only source — small, mergeable diffs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;nbdime&lt;/code&gt;.&lt;/strong&gt; Content-aware diff/merge: &lt;code&gt;nbdiff&lt;/code&gt; for the terminal, &lt;code&gt;nbdiff-web&lt;/code&gt; for a rendered side-by-side, and &lt;code&gt;nbdime config-git --enable&lt;/code&gt; to make &lt;code&gt;git diff&lt;/code&gt; / &lt;code&gt;git merge&lt;/code&gt; understand notebooks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;jupytext pairing.&lt;/strong&gt; Pair each &lt;code&gt;.ipynb&lt;/code&gt; with a synced &lt;code&gt;.py&lt;/code&gt; "percent" script so pull requests review clean Python while the notebook stays runnable — a common belt-and-suspenders setup.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feb37ev1rnsluyp0o9tue.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feb37ev1rnsluyp0o9tue.jpeg" alt="Iconographic diagram — an Airflow PapermillOperator task running a notebook on a schedule with a per-run output path, alongside nbdime and nbstripout cleaning noisy notebook JSON diffs for version control." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — a scheduled PapermillOperator task
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The standard way to schedule a notebook is a single Airflow task built from &lt;code&gt;PapermillOperator&lt;/code&gt;, templated so each daily run injects that day's logical date and writes a dated artifact. The DAG owns cadence and retries; the operator owns the Papermill call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Define a daily Airflow task that runs &lt;code&gt;etl.ipynb&lt;/code&gt;, passing the run's logical date as &lt;code&gt;run_date&lt;/code&gt; and writing the executed notebook to a dated path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;field&lt;/th&gt;
&lt;th&gt;value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;schedule&lt;/td&gt;
&lt;td&gt;daily (&lt;code&gt;@daily&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;input notebook&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/nbs/etl.ipynb&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;output notebook&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/runs/etl-{{ ds }}.ipynb&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;parameter&lt;/td&gt;
&lt;td&gt;&lt;code&gt;run_date = {{ ds }}&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;airflow&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DAG&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;airflow.providers.papermill.operators.papermill&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;PapermillOperator&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pendulum&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nc"&gt;DAG&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;dag_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;etl_notebook&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;schedule&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;@daily&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;start_date&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;pendulum&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2026&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tz&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;UTC&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;catchup&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;dag&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;run_etl&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;PapermillOperator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;task_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_etl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;input_nb&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/nbs/etl.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;output_nb&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/runs/etl-{{ ds }}.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="c1"&gt;# dated artifact per run
&lt;/span&gt;        &lt;span class="n"&gt;parameters&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{ ds }}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;        &lt;span class="c1"&gt;# logical date, not now()
&lt;/span&gt;    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Airflow schedules &lt;code&gt;etl_notebook&lt;/code&gt; daily. For each run it renders the Jinja templates: &lt;code&gt;{{ ds }}&lt;/code&gt; becomes that interval's logical date (e.g. &lt;code&gt;2026-09-15&lt;/code&gt;). &lt;code&gt;PapermillOperator&lt;/code&gt; then calls &lt;code&gt;execute_notebook("/nbs/etl.ipynb", "/runs/etl-2026-09-15.ipynb", parameters={"run_date": "2026-09-15"})&lt;/code&gt;. Because the date comes from the scheduler's logical date, re-running the task reproduces the same output; retries are safe. Each day leaves its own dated artifact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;run interval&lt;/th&gt;
&lt;th&gt;output notebook&lt;/th&gt;
&lt;th&gt;injected &lt;code&gt;run_date&lt;/code&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-15&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/runs/etl-2026-09-15.ipynb&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;2026-09-15&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-16&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/runs/etl-2026-09-16.ipynb&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;2026-09-16&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Parametrize by the orchestrator's logical run date, never wall-clock &lt;code&gt;now()&lt;/code&gt; — that is what makes a scheduled notebook idempotent and backfillable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Papermill interview question on notebooks in version control
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Reviewers refuse to approve notebook PRs because "the diffs are unreadable" — a two-line code change shows thousands of changed lines. The team still wants to keep working in notebooks. What do you set up so notebook changes review cleanly without abandoning Jupyter?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using nbstripout and nbdime as git integrations
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;nbstripout nbdime
nbstripout &lt;span class="nt"&gt;--install&lt;/span&gt;                       &lt;span class="c"&gt;# git clean filter: strip outputs on commit&lt;/span&gt;
nbdime config-git &lt;span class="nt"&gt;--enable&lt;/span&gt; &lt;span class="nt"&gt;--global&lt;/span&gt;        &lt;span class="c"&gt;# make git diff/merge notebook-aware&lt;/span&gt;
nbdiff-web notebooks/etl.ipynb             &lt;span class="c"&gt;# rendered side-by-side diff in the browser&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;tool&lt;/th&gt;
&lt;th&gt;effect on the repo / review&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;nbstripout --install&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;commits drop outputs + &lt;code&gt;execution_count&lt;/code&gt; → tiny diffs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;&lt;code&gt;nbdime config-git --enable&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;git diff&lt;/code&gt; shows cell-level source changes, not JSON&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;nbdiff-web&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;renders a two-column diff a reviewer can actually read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;merge conflict&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;nbmerge&lt;/code&gt; resolves per-cell instead of per-JSON-line&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;nbstripout&lt;/code&gt; installs a git &lt;em&gt;clean&lt;/em&gt; filter that removes outputs and execution counts before content is staged, so the stored notebook is mostly source — the base64 image blobs that bloated the diff are gone.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;nbdime config-git --enable&lt;/code&gt; registers nbdime as git's diff and merge driver for &lt;code&gt;.ipynb&lt;/code&gt;, so &lt;code&gt;git diff&lt;/code&gt; compares notebooks by cell content rather than raw JSON lines.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;nbdiff-web&lt;/code&gt; gives reviewers a rendered, side-by-side view where a two-line change looks like a two-line change.&lt;/li&gt;
&lt;li&gt;When two branches edit the same notebook, &lt;code&gt;nbmerge&lt;/code&gt; reconciles cell-by-cell, avoiding the unresolvable JSON conflicts that make notebook collaboration painful.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;before&lt;/th&gt;
&lt;th&gt;after&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2-line change → thousands of diff lines&lt;/td&gt;
&lt;td&gt;2-line change → 2 diff lines&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;unmergeable JSON conflicts&lt;/td&gt;
&lt;td&gt;cell-level &lt;code&gt;nbmerge&lt;/code&gt; resolution&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Strip outputs at the boundary&lt;/strong&gt;&lt;/strong&gt; — nbstripout removes the volatile, huge parts (outputs, counts) at commit time, so version control tracks intent, not render noise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Content-aware diff&lt;/strong&gt;&lt;/strong&gt; — nbdime compares the notebook's cell structure rather than its serialized JSON, turning an unreadable blob diff into a code review.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Git integration&lt;/strong&gt;&lt;/strong&gt; — wiring both tools into git means reviewers get clean diffs with no change to how authors work in Jupyter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Mergeable notebooks&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;nbmerge&lt;/code&gt; resolves per cell, removing the "notebooks can't be collaborated on" objection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — filters and diffs are O(notebook size) at commit/review time; the throughput cost is trivial next to the review friction they remove.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Scheduling&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — scheduling&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Scheduled-run and idempotency problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/scheduling" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Pipelines&lt;/span&gt;
&lt;span&gt;Topic — pipelines&lt;/span&gt;
&lt;strong&gt;Orchestration and DAG-design problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — Papermill recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Minimal execute.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;papermill&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pm&lt;/span&gt;
&lt;span class="n"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute_notebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;in.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;out.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;parameters&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2026-09-15&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;CLI parametrized run.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;papermill &lt;span class="k"&gt;in&lt;/span&gt;.ipynb out.ipynb &lt;span class="nt"&gt;-p&lt;/span&gt; region eu &lt;span class="nt"&gt;-p&lt;/span&gt; limit 500 &lt;span class="nt"&gt;-r&lt;/span&gt; zip 02139 &lt;span class="nt"&gt;-k&lt;/span&gt; python3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;YAML params file.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;papermill &lt;span class="k"&gt;in&lt;/span&gt;.ipynb out.ipynb &lt;span class="nt"&gt;-f&lt;/span&gt; params.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;where &lt;code&gt;params.yaml&lt;/code&gt; contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;region&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;eu&lt;/span&gt;
&lt;span class="na"&gt;skus&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;a&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;b&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;c&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Glue a value and read it back.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;scrapbook&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;sb&lt;/span&gt;                &lt;span class="c1"&gt;# inside the notebook
&lt;/span&gt;&lt;span class="n"&gt;sb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;glue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;row_count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="n"&gt;nb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_notebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;out.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# downstream: read it back
&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;nb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;scraps&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;row_count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Airflow PapermillOperator task.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;airflow.providers.papermill.operators.papermill&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;PapermillOperator&lt;/span&gt;
&lt;span class="nc"&gt;PapermillOperator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;task_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_etl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;input_nb&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/nbs/etl.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;output_nb&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/runs/etl-{{ ds }}.ipynb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;parameters&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{ ds }}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Clean notebook diffs in git.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nbstripout &lt;span class="nt"&gt;--install&lt;/span&gt;                 &lt;span class="c"&gt;# strip outputs on commit&lt;/span&gt;
nbdime config-git &lt;span class="nt"&gt;--enable&lt;/span&gt; &lt;span class="nt"&gt;--global&lt;/span&gt;  &lt;span class="c"&gt;# notebook-aware diff/merge&lt;/span&gt;
jupyter nbconvert &lt;span class="nt"&gt;--to&lt;/span&gt; html &lt;span class="nt"&gt;--no-input&lt;/span&gt; out.ipynb   &lt;span class="c"&gt;# code-free report&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Choosing the pattern.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Run the same notebook per date / region&lt;/td&gt;
&lt;td&gt;Papermill parameters + dated output path&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Get a scalar / frame back to the caller&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;scrapbook.glue&lt;/code&gt; + &lt;code&gt;read_notebook&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trigger on a schedule with retries&lt;/td&gt;
&lt;td&gt;Airflow &lt;code&gt;PapermillOperator&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Share outputs without code&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;nbconvert --no-input&lt;/code&gt; (HTML/PDF)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Keep notebook PRs reviewable&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;nbstripout&lt;/code&gt; + &lt;code&gt;nbdime&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is Papermill?
&lt;/h3&gt;

&lt;p&gt;Papermill is an open-source tool (a Python library and a CLI) for &lt;strong&gt;parametrizing and executing Jupyter notebooks&lt;/strong&gt;. You pass parameters into a notebook without editing its source, Papermill runs every cell against a kernel, and it writes a new &lt;em&gt;output notebook&lt;/em&gt; containing the code, the injected parameters, and every output the run produced. It does not schedule or transform — it executes a notebook and hands you the executed artifact.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I parametrize a Jupyter notebook with Papermill?
&lt;/h3&gt;

&lt;p&gt;Tag one cell with the cell tag &lt;code&gt;parameters&lt;/code&gt; and put default variable assignments in it. At run time Papermill injects a new &lt;code&gt;injected-parameters&lt;/code&gt; cell immediately after the tagged cell, overriding those defaults with the values you pass via &lt;code&gt;pm.execute_notebook(..., parameters={...})&lt;/code&gt; or the CLI &lt;code&gt;-p&lt;/code&gt;/&lt;code&gt;-r&lt;/code&gt;/&lt;code&gt;-y&lt;/code&gt;/&lt;code&gt;-f&lt;/code&gt; flags. The defaults keep the notebook runnable by hand; the injected values win because they run last.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the output notebook in Papermill?
&lt;/h3&gt;

&lt;p&gt;It is the executed artifact Papermill writes for each run — the input notebook with every cell run, all outputs captured inline (tables, logs, charts), per-cell timings recorded, and the injected-parameters cell visible. The input template is never modified. If a cell raises, Papermill still saves the output notebook with the traceback in place, so a failed run remains fully debuggable.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I get values back out of a Papermill notebook?
&lt;/h3&gt;

&lt;p&gt;Use the companion library &lt;strong&gt;scrapbook&lt;/strong&gt; (Papermill's old &lt;code&gt;record&lt;/code&gt;/&lt;code&gt;read_notebook&lt;/code&gt; API is deprecated). Inside the notebook call &lt;code&gt;sb.glue("name", value)&lt;/code&gt; to store a value in the output notebook; downstream, call &lt;code&gt;sb.read_notebook(path).scraps["name"].data&lt;/code&gt; to read it back with its type intact. For a batch of runs, &lt;code&gt;sb.read_notebooks(dir)&lt;/code&gt; assembles one row per notebook so you can compare metrics across a parametrized fan-out.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I schedule Papermill notebooks?
&lt;/h3&gt;

&lt;p&gt;Papermill has no scheduler, so you trigger it from an orchestrator. In Airflow, the &lt;code&gt;PapermillOperator&lt;/code&gt; (from &lt;code&gt;airflow.providers.papermill&lt;/code&gt;) wraps &lt;code&gt;execute_notebook&lt;/code&gt;: give it &lt;code&gt;input_nb&lt;/code&gt;, a templated &lt;code&gt;output_nb&lt;/code&gt;, and &lt;code&gt;parameters&lt;/code&gt;, and template the run's logical date (&lt;code&gt;{{ ds }}&lt;/code&gt;) instead of &lt;code&gt;now()&lt;/code&gt; so re-runs are idempotent. Retries, alerts, and SLAs live on the Airflow task; Papermill just runs the notebook and fails loudly if a cell errors.&lt;/p&gt;

&lt;h3&gt;
  
  
  When are notebooks-as-pipelines a bad idea?
&lt;/h3&gt;

&lt;p&gt;When the logic is complex, shared, or heavily branched. Notebooks encourage hidden execution state, they are awkward to unit-test, and their JSON format makes multi-author collaboration and code review painful. Papermill mitigates the run-order problem by always executing top-to-bottom, but for reusable libraries, fan-out DAGs, or team-owned pipeline logic, extract the code into tested &lt;code&gt;.py&lt;/code&gt; modules and call them from tasks — keep notebooks for the reporting and exploratory jobs where the rendered artifact is the deliverable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every Papermill idea above, from the parameters cell tag to the executed output artifact, scrapbook glue, and the scheduled PapermillOperator, maps to a hands-on practice room where you design the pipeline against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you make this scheduled notebook idempotent?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice pipeline problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/scheduling" rel="noopener noreferrer"&gt;Scheduling drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Streamlit for Data Engineers: Internal Data Apps &amp; Pipeline Dashboards</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Thu, 24 Sep 2026 17:40:15 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/streamlit-for-data-engineers-internal-data-apps-pipeline-dashboards-483b</link>
      <guid>https://dev.to/gowthampotureddi/streamlit-for-data-engineers-internal-data-apps-pipeline-dashboards-483b</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;streamlit for data engineers&lt;/code&gt;&lt;/strong&gt; answers a very specific, very common need: you have a warehouse full of pipeline metadata, freshness timestamps, and row counts, and someone — an analyst, an on-call engineer, your own future self at 3 a.m. — needs to &lt;em&gt;see&lt;/em&gt; it and &lt;em&gt;act&lt;/em&gt; on it without you standing up a React frontend and a Flask API. Streamlit turns a single Python script into a web app. You write &lt;code&gt;import streamlit as st&lt;/code&gt;, add a few &lt;code&gt;st.&lt;/code&gt; calls that print DataFrames and draw charts, run &lt;code&gt;streamlit run app.py&lt;/code&gt;, and you have an internal tool on a URL. No HTML, no JavaScript, no callback wiring, no template engine.&lt;/p&gt;

&lt;p&gt;That shape is a real departure from the two things data engineers reached for before it: a Jupyter notebook that only you can run and that no non-technical teammate will ever open, or a "proper" web app that costs a week of frontend work you do not want to own. Streamlit lives in between — it is code-first like the notebook and shareable like the web app. This guide walks the four ideas an interviewer or a code reviewer will actually probe when you put Streamlit in your stack — the top-to-bottom rerun model with widgets and &lt;code&gt;session_state&lt;/code&gt;, the two caching decorators, displaying data and connecting to a warehouse, and assembling all of it into a pipeline-health / SLA dashboard — and pairs each with a Solution-Tail interview answer: code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd0bgyotpq8txsudd53v8.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd0bgyotpq8txsudd53v8.jpeg" alt="PipeCode blog header for Streamlit for data engineers — bold white headline 'Streamlit for Data Engineers' with subtitle 'Internal Apps · Pipeline Dashboards' and a stylised script-to-dashboard scene on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the DataFrame shaping that feeds every widget on the &lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;data-analysis practice library →&lt;/a&gt;, sharpen the transforms behind your tables on the &lt;a href="https://pipecode.ai/explore/practice/topic/dataframe-basics" rel="noopener noreferrer"&gt;dataframe-basics practice set →&lt;/a&gt;, and rehearse the freshness-and-breach logic your dashboard renders on the &lt;a href="https://pipecode.ai/explore/practice/topic/sla-monitoring" rel="noopener noreferrer"&gt;sla-monitoring practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why Streamlit fits data engineering in 2026&lt;/li&gt;
&lt;li&gt;The rerun model, widgets &amp;amp; session_state&lt;/li&gt;
&lt;li&gt;Caching — st.cache_data vs st.cache_resource&lt;/li&gt;
&lt;li&gt;Displaying data &amp;amp; connecting to warehouses&lt;/li&gt;
&lt;li&gt;Building a pipeline-health / SLA dashboard&lt;/li&gt;
&lt;li&gt;Cheat sheet — Streamlit recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why Streamlit fits data engineering in 2026
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Streamlit is a Python script that renders as a web app — that one fact decides where it fits and how it behaves
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;a Streamlit app is an ordinary Python script that the framework re-executes top to bottom every time the user interacts with it, turning &lt;code&gt;st.&lt;/code&gt; calls into UI&lt;/strong&gt;. Everything surprising about Streamlit — why a variable resets, why caching matters so much, why you need &lt;code&gt;session_state&lt;/code&gt; — falls out of that single sentence. There is no component tree you register, no event loop you write, no separate frontend build. You write the script the way you would write an analysis, and Streamlit renders it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What Streamlit does — and deliberately does not — do.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;UI from Python.&lt;/strong&gt; &lt;code&gt;st.dataframe(df)&lt;/code&gt;, &lt;code&gt;st.line_chart(df)&lt;/code&gt;, &lt;code&gt;st.button("Run")&lt;/code&gt;, &lt;code&gt;st.metric("Rows", 1200)&lt;/code&gt; — each call emits a widget. You never touch HTML or CSS for the common cases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;State and rerun handled for you.&lt;/strong&gt; The framework owns the render loop: interact with a widget and the whole script reruns. You do not manage a DOM or diff a virtual tree.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not a BI product, not a heavy web framework.&lt;/strong&gt; Streamlit is not trying to be Tableau (no semantic layer, no governed metrics) and not trying to be Django (no ORM, no auth framework, no routing beyond simple multipage). It is the fastest path from &lt;em&gt;"I have a DataFrame"&lt;/em&gt; to &lt;em&gt;"my team can click on it."&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Where Streamlit sits against the alternatives data engineers weigh.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;vs a Jupyter notebook.&lt;/strong&gt; A notebook is for &lt;em&gt;you&lt;/em&gt; exploring; a Streamlit app is for &lt;em&gt;others&lt;/em&gt; using. The notebook has hidden execution order and stale cells; Streamlit reruns cleanly every time, so what the viewer sees always matches the current inputs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vs Flask / FastAPI + a JS frontend.&lt;/strong&gt; A hand-rolled web app gives you total control and costs days of frontend work you must maintain. Streamlit trades that control for a Python-only script you finish in an afternoon — the right trade for an internal tool with ten users, the wrong one for a public product.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vs Dash / Panel.&lt;/strong&gt; Dash is callback-oriented (you wire inputs to outputs explicitly); Streamlit is rerun-oriented (no callbacks needed for the basics). Streamlit is faster to write; Dash gives finer-grained control when you truly need it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vs a BI dashboard (Looker, Metabase).&lt;/strong&gt; BI tools own governed metrics and scheduled reports; Streamlit owns &lt;em&gt;interactive tools that do things&lt;/em&gt; — trigger a backfill, edit a config, kick a re-run — which a read-only BI tile cannot.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you say &lt;strong&gt;"the whole script reruns on every interaction"&lt;/strong&gt; in the first sentence? — this is the senior signal; everything else is a corollary.&lt;/li&gt;
&lt;li&gt;Do you reach for &lt;strong&gt;&lt;code&gt;st.cache_data&lt;/code&gt; / &lt;code&gt;st.cache_resource&lt;/code&gt;&lt;/strong&gt; unprompted when you mention loading data, because you know the rerun would otherwise re-query every time? — required framing.&lt;/li&gt;
&lt;li&gt;Do you know that a plain Python variable &lt;strong&gt;does not survive a rerun&lt;/strong&gt;, and that &lt;code&gt;st.session_state&lt;/code&gt; is the fix? — the number-one beginner gap.&lt;/li&gt;
&lt;li&gt;Do you place Streamlit as &lt;strong&gt;"internal tools and dashboards, not a public product"&lt;/strong&gt; rather than as a Flask replacement for everything? — placement maturity.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — six lines that become a shareable app
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The canonical Streamlit "hello world" reads a DataFrame and renders it as a title, an interactive table, and a chart. It looks trivial, and that is the point: the same six lines that render a toy DataFrame scale unchanged to a live warehouse query, because Streamlit only ever sees a DataFrame and a few &lt;code&gt;st.&lt;/code&gt; calls. There is no server code, no route, no template — the script &lt;em&gt;is&lt;/em&gt; the app.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Turn a two-row pipeline-runs DataFrame into a titled app with an interactive table and a bar chart, runnable with one command.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;pipeline&lt;/th&gt;
&lt;th&gt;rows_loaded&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;orders_el&lt;/td&gt;
&lt;td&gt;4200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;events_el&lt;/td&gt;
&lt;td&gt;9100&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;streamlit&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;

&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;title&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Pipeline runs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pipeline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;orders_el&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;events_el&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rows_loaded&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;4200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;9100&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dataframe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                       &lt;span class="c1"&gt;# interactive, sortable table
&lt;/span&gt;&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;bar_chart&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pipeline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rows_loaded&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;import streamlit as st&lt;/code&gt; pulls in the whole UI toolkit. &lt;code&gt;st.title(...)&lt;/code&gt; emits an &lt;code&gt;&amp;lt;h1&amp;gt;&lt;/code&gt; at the top of the page. Building &lt;code&gt;df&lt;/code&gt; is ordinary pandas — Streamlit has no opinion about how you get the DataFrame. &lt;code&gt;st.dataframe(df)&lt;/code&gt; renders a sortable, scrollable, searchable grid; &lt;code&gt;st.bar_chart(...)&lt;/code&gt; draws a chart from the same frame with no plotting library imported. You launch it with &lt;code&gt;streamlit run app.py&lt;/code&gt;, which starts a local server and opens the browser; every save auto-reloads.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;the app shows&lt;/th&gt;
&lt;th&gt;rendered from&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;page title "Pipeline runs"&lt;/td&gt;
&lt;td&gt;&lt;code&gt;st.title(...)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;an interactive 2-row table&lt;/td&gt;
&lt;td&gt;&lt;code&gt;st.dataframe(df)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a bar chart, pipeline vs rows_loaded&lt;/td&gt;
&lt;td&gt;&lt;code&gt;st.bar_chart(...)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a live URL on &lt;code&gt;localhost:8501&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;streamlit run app.py&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If your logic ends in a DataFrame, a number, or a chart, Streamlit can display it — the toy example and the production dashboard differ only in where the DataFrame comes from, never in the &lt;code&gt;st.&lt;/code&gt; calls that render it.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. The rerun model, widgets &amp;amp; session_state
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The whole script reruns top-to-bottom on every interaction — widgets return values, &lt;code&gt;session_state&lt;/code&gt; is what survives
&lt;/h3&gt;

&lt;p&gt;If you internalise one thing about Streamlit, make it this: &lt;strong&gt;there is no event loop and no callback graph; every widget interaction re-executes your entire script from the first line to the last&lt;/strong&gt;. A button click, a slider drag, a dropdown change — each schedules a fresh top-to-bottom run. Widgets are not "handlers you attach"; they are function calls that &lt;em&gt;return their current value&lt;/em&gt; on each run. This is why a normal variable resets and why state needs a special home.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The rerun loop, precisely.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Interaction → rerun.&lt;/strong&gt; The user moves a slider; Streamlit re-runs &lt;code&gt;app.py&lt;/code&gt; from the top. Your code executes again, the slider call returns the new value, and the page is redrawn from scratch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Widgets return values, not events.&lt;/strong&gt; &lt;code&gt;value = st.slider("Days", 1, 30, 7)&lt;/code&gt; returns &lt;code&gt;7&lt;/code&gt; on the first run and whatever the user picked on later runs. You use the return value inline — no &lt;code&gt;onChange&lt;/code&gt; handler.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plain variables do not persist.&lt;/strong&gt; Any &lt;code&gt;x = 0&lt;/code&gt; at the top is re-initialised to &lt;code&gt;0&lt;/code&gt; on every rerun. If you increment it on a click, the increment is wiped by the next rerun. This is the single most common beginner bug.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;st.session_state&lt;/code&gt; — the store that survives reruns.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A dict scoped to the browser session.&lt;/strong&gt; &lt;code&gt;st.session_state&lt;/code&gt; is a dictionary-like object that persists across reruns &lt;em&gt;for one user session&lt;/em&gt;. Write &lt;code&gt;st.session_state["count"] = 0&lt;/code&gt; once, read and mutate it on later runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Widgets can bind to it via &lt;code&gt;key&lt;/code&gt;.&lt;/strong&gt; Give a widget &lt;code&gt;key="days"&lt;/code&gt; and its value is mirrored at &lt;code&gt;st.session_state["days"]&lt;/code&gt;; reading the key elsewhere gives the live value without threading it through function arguments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Initialise defensively.&lt;/strong&gt; Because the script reruns, guard initialisation with &lt;code&gt;if "count" not in st.session_state:&lt;/code&gt; so you seed the value once and never clobber it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Callbacks and the button gotcha.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Callbacks run &lt;em&gt;before&lt;/em&gt; the rerun.&lt;/strong&gt; &lt;code&gt;st.button("Add", on_click=fn)&lt;/code&gt; runs &lt;code&gt;fn&lt;/code&gt; first, then reruns the script. Callbacks (&lt;code&gt;on_click&lt;/code&gt;, &lt;code&gt;on_change&lt;/code&gt;) are where you mutate &lt;code&gt;session_state&lt;/code&gt; cleanly, and they receive &lt;code&gt;args&lt;/code&gt; / &lt;code&gt;kwargs&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.button&lt;/code&gt; is transient.&lt;/strong&gt; &lt;code&gt;st.button(...)&lt;/code&gt; returns &lt;code&gt;True&lt;/code&gt; only on the single rerun immediately following the click, then &lt;code&gt;False&lt;/code&gt; again. So &lt;code&gt;if st.button("Go"): x = compute()&lt;/code&gt; recomputes once — do not expect the &lt;code&gt;True&lt;/code&gt; to "stick." For durable toggles use &lt;code&gt;st.session_state&lt;/code&gt; or &lt;code&gt;st.checkbox&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg43lbvqyhdvoe5vs391s.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg43lbvqyhdvoe5vs391s.jpeg" alt="Iconographic Streamlit rerun-model diagram — a widget interaction triggering a full top-to-bottom re-execution of the script, widgets returning their current values, and a session_state store persisting keys across reruns." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — a counter that actually counts
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The classic demonstration of the rerun model is a click counter, because the naive version is broken in a way that teaches the whole model. A plain variable is reset by the rerun; &lt;code&gt;session_state&lt;/code&gt; plus a callback is the fix. Getting this right is the difference between an app that "forgets" and one that holds state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Build a button that increments a counter and shows the running total, surviving every rerun.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; Three clicks of the "Add one" button.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;streamlit&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;     &lt;span class="c1"&gt;# seed once, never on later reruns
&lt;/span&gt;    &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;increment&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;                        &lt;span class="c1"&gt;# callback runs BEFORE the rerun
&lt;/span&gt;    &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;

&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;button&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Add one&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;on_click&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;increment&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Count:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; On the first run, &lt;code&gt;count&lt;/code&gt; is absent, so it is seeded to &lt;code&gt;0&lt;/code&gt;. Clicking the button fires &lt;code&gt;increment&lt;/code&gt; &lt;em&gt;before&lt;/em&gt; the rerun, so &lt;code&gt;count&lt;/code&gt; becomes &lt;code&gt;1&lt;/code&gt;; then the script reruns, the &lt;code&gt;if&lt;/code&gt; guard is skipped (key exists), and &lt;code&gt;st.write&lt;/code&gt; prints &lt;code&gt;1&lt;/code&gt;. Each subsequent click repeats: callback mutates state, script reruns, the guard leaves the existing value alone. Had we written &lt;code&gt;count = 0&lt;/code&gt; as a plain variable at the top, every rerun would reset it to &lt;code&gt;0&lt;/code&gt; and the display would never exceed &lt;code&gt;1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;click #&lt;/th&gt;
&lt;th&gt;value before callback&lt;/th&gt;
&lt;th&gt;value shown&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Anything that must outlive a rerun lives in &lt;code&gt;st.session_state&lt;/code&gt;; mutate it in a callback, and guard its initialisation with an &lt;code&gt;if key not in st.session_state&lt;/code&gt; check.&lt;/p&gt;

&lt;h3&gt;
  
  
  Streamlit interview question on the rerun model
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A candidate writes &lt;code&gt;clicks = 0&lt;/code&gt; at the top of the script and does &lt;code&gt;if st.button("Click"): clicks += 1; st.write(clicks)&lt;/code&gt;. The display never goes past &lt;code&gt;1&lt;/code&gt;. Explain exactly why, and fix it so the count accumulates.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using session_state seeded once with a callback
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;streamlit&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;

&lt;span class="c1"&gt;# BROKEN: `clicks` is a plain variable, reset to 0 on every rerun
# clicks = 0
# if st.button("Click"):
#     clicks += 1
# st.write(clicks)          # always 0 or 1
&lt;/span&gt;
&lt;span class="c1"&gt;# FIXED: state lives in session_state, mutated in a callback
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;clicks&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;clicks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;bump&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;clicks&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;

&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;button&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Click&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;on_click&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;bump&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;clicks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;event&lt;/th&gt;
&lt;th&gt;rerun #&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;clicks&lt;/code&gt; at top&lt;/th&gt;
&lt;th&gt;after handler&lt;/th&gt;
&lt;th&gt;displayed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;initial load&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;seeded 0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;click&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0 (state)&lt;/td&gt;
&lt;td&gt;bump → 1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;click&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;1 (state)&lt;/td&gt;
&lt;td&gt;bump → 2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;click&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;2 (state)&lt;/td&gt;
&lt;td&gt;bump → 3&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;In the broken version, &lt;code&gt;clicks = 0&lt;/code&gt; executes on &lt;strong&gt;every&lt;/strong&gt; rerun, so the pre-increment value is always &lt;code&gt;0&lt;/code&gt;; the click makes it &lt;code&gt;1&lt;/code&gt;, and the next rerun resets it. The counter can never exceed &lt;code&gt;1&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;st.button(...)&lt;/code&gt; returns &lt;code&gt;True&lt;/code&gt; only on the rerun right after the click, so the &lt;code&gt;+= 1&lt;/code&gt; fires at most once per click — but against a variable that was just reset.&lt;/li&gt;
&lt;li&gt;The fix seeds &lt;code&gt;clicks&lt;/code&gt; in &lt;code&gt;session_state&lt;/code&gt; &lt;strong&gt;once&lt;/strong&gt;, guarded by the &lt;code&gt;not in&lt;/code&gt; check, so it is not clobbered on reruns.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;on_click=bump&lt;/code&gt; callback mutates the persisted value &lt;em&gt;before&lt;/em&gt; the redraw, so &lt;code&gt;st.write&lt;/code&gt; reads the accumulated total.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;after 3 clicks&lt;/th&gt;
&lt;th&gt;broken version&lt;/th&gt;
&lt;th&gt;fixed version&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;displayed count&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Top-to-bottom rerun&lt;/strong&gt;&lt;/strong&gt; — because the entire script re-executes on each interaction, any assignment at module scope is re-run and therefore re-initialised; state cannot live in a local variable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;session_state persistence&lt;/strong&gt;&lt;/strong&gt; — a dict scoped to the session survives reruns, so a value written once is readable and mutable on every later run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Guarded initialisation&lt;/strong&gt;&lt;/strong&gt; — the &lt;code&gt;if key not in st.session_state&lt;/code&gt; seed runs on the first rerun only, protecting the accumulated value from being reset.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Callback ordering&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;on_click&lt;/code&gt; runs before the rerun, so the mutation is visible to the redraw in the same cycle, avoiding an off-by-one lag.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the rerun is O(script length) each interaction; state access is O(1), which is why heavy work must be cached rather than repeated.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Analysis&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — data-analysis&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Interactive data-analysis app problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;DataFrames&lt;/span&gt;
&lt;span&gt;Topic — dataframe-basics&lt;/span&gt;
&lt;strong&gt;DataFrame state-and-filter problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/dataframe-basics" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. Caching — st.cache_data vs st.cache_resource
&lt;/h2&gt;
&lt;h3&gt;
  
  
  The rerun re-runs everything, so caching is not optional — pick &lt;code&gt;st.cache_data&lt;/code&gt; for values and &lt;code&gt;st.cache_resource&lt;/code&gt; for connections
&lt;/h3&gt;

&lt;p&gt;Because the whole script reruns on every click, an uncached &lt;code&gt;pd.read_sql(...)&lt;/code&gt; re-queries the warehouse every time a user drags a slider — unusable. Streamlit's answer is two decorators, and the interview question is always "which one and why." Say it in one breath: &lt;strong&gt;&lt;code&gt;st.cache_data&lt;/code&gt; memoizes the &lt;em&gt;return value&lt;/em&gt; and hands each caller a fresh copy; &lt;code&gt;st.cache_resource&lt;/code&gt; caches a &lt;em&gt;single global object&lt;/em&gt; and hands every caller the same instance&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;st.cache_data&lt;/code&gt; — for data you compute or fetch.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Memoizes serializable results.&lt;/strong&gt; Wrap a function that returns a DataFrame, a dict, a list, an API response — anything picklable. On a cache hit Streamlit skips the body and returns the stored value.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keyed on the function's inputs.&lt;/strong&gt; Streamlit hashes the arguments (and the function's code); same args → cache hit, different args → recompute. So &lt;code&gt;load(start, end)&lt;/code&gt; caches per date range automatically.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Returns a copy every time.&lt;/strong&gt; Each caller gets its own copy of the result, so one user mutating the DataFrame cannot corrupt another user's cached copy. That safety is the whole reason it is separate from &lt;code&gt;cache_resource&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;st.cache_resource&lt;/code&gt; — for global singletons.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Caches the object itself, not a copy.&lt;/strong&gt; A database connection, a SQLAlchemy engine, an ML model, a client handle — things that are expensive to create and meant to be shared. Every caller receives the &lt;em&gt;same&lt;/em&gt; live object.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not serialized, not copied.&lt;/strong&gt; Because it is shared, it must be safe for concurrent use; a connection pool or a thread-safe client is the right kind of thing to put here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One per app, across sessions.&lt;/strong&gt; Unlike &lt;code&gt;session_state&lt;/code&gt; (per session), a &lt;code&gt;cache_resource&lt;/code&gt; object is shared across all users and reruns until the app restarts or you clear it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The knobs and the failure modes.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ttl&lt;/code&gt;.&lt;/strong&gt; Expire entries after N seconds (&lt;code&gt;ttl=600&lt;/code&gt;) so a dashboard shows data at most ten minutes stale without a manual refresh.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;max_entries&lt;/code&gt;.&lt;/strong&gt; Cap the number of cached results (e.g. &lt;code&gt;max_entries=50&lt;/code&gt;) to bound memory when the argument space is large.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;show_spinner&lt;/code&gt; and &lt;code&gt;.clear()&lt;/code&gt;.&lt;/strong&gt; Toggle the "Running..." spinner, and call &lt;code&gt;load.clear()&lt;/code&gt; to evict programmatically (e.g. after a backfill writes new rows).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The classic mistake:&lt;/strong&gt; putting a database connection in &lt;code&gt;st.cache_data&lt;/code&gt; (it tries to pickle the connection and fails or misbehaves), or putting a DataFrame you mutate in &lt;code&gt;st.cache_resource&lt;/code&gt; (mutations leak across users because there is no copy).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnfo88d9w0uu2z58r8vc9.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnfo88d9w0uu2z58r8vc9.jpeg" alt="Iconographic Streamlit caching diagram — st.cache_data returning a fresh copy of a DataFrame keyed on function arguments, versus st.cache_resource returning one shared connection singleton, with a TTL dial." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — cache the query result, cache the engine once
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday pattern pairs the two decorators: &lt;code&gt;cache_resource&lt;/code&gt; opens the SQLAlchemy engine one time and shares it; &lt;code&gt;cache_data&lt;/code&gt; runs the query and memoizes the resulting DataFrame per set of parameters. Together they turn a per-rerun round trip into a per-argument-set round trip.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Write a helper that opens a Postgres engine once and a query function that caches its DataFrame for five minutes, keyed on the day range.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; Two reruns with the same &lt;code&gt;days=7&lt;/code&gt;, then one with &lt;code&gt;days=30&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;streamlit&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sqlalchemy&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;create_engine&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;

&lt;span class="nd"&gt;@st.cache_resource&lt;/span&gt;                      &lt;span class="c1"&gt;# one engine, shared across all reruns/users
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_engine&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;create_engine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;secrets&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;db_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="nd"&gt;@st.cache_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ttl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                 &lt;span class="c1"&gt;# memoize the DataFrame per `days`, 5-min TTL
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;load_runs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;days&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;sql&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT pipeline, status, ended_at FROM runs &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
               &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;WHERE ended_at &amp;gt; now() - make_interval(days =&amp;gt; :d)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;get_engine&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_sql&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sql&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;d&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;days&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="n"&gt;days&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Look-back (days)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dataframe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;load_runs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;days&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The first rerun calls &lt;code&gt;get_engine()&lt;/code&gt;, which builds the engine and caches the object; &lt;code&gt;load_runs(7)&lt;/code&gt; runs the SQL and caches the DataFrame under key &lt;code&gt;days=7&lt;/code&gt;. Dragging the slider and releasing at &lt;code&gt;7&lt;/code&gt; again reruns the script, but &lt;code&gt;get_engine()&lt;/code&gt; returns the cached engine and &lt;code&gt;load_runs(7)&lt;/code&gt; returns the cached DataFrame — zero database work. Moving the slider to &lt;code&gt;30&lt;/code&gt; is a new argument, so &lt;code&gt;load_runs(30)&lt;/code&gt; misses the cache and queries once, then caches that too. After five minutes the &lt;code&gt;days=7&lt;/code&gt; entry expires and the next access re-queries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;rerun&lt;/th&gt;
&lt;th&gt;engine built?&lt;/th&gt;
&lt;th&gt;query run?&lt;/th&gt;
&lt;th&gt;why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 (days=7)&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;cold cache&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 (days=7)&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;both cache hits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 (days=30)&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;new arg → data miss&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4 (days=7, &amp;gt;5 min later)&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;ttl expired&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If the function returns &lt;em&gt;data&lt;/em&gt;, decorate with &lt;code&gt;st.cache_data&lt;/code&gt;; if it returns a &lt;em&gt;thing you open once and reuse&lt;/em&gt; (a connection, an engine, a client, a model), decorate with &lt;code&gt;st.cache_resource&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Streamlit interview question on choosing a cache decorator
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; An engineer decorates their &lt;code&gt;get_connection()&lt;/code&gt; function with &lt;code&gt;@st.cache_data&lt;/code&gt; and their &lt;code&gt;load_orders()&lt;/code&gt; DataFrame function with &lt;code&gt;@st.cache_resource&lt;/code&gt;. The app is slow and occasionally shows another user's filtered data. Which decorators are swapped, and what breaks in each case?&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using cache_resource for the connection and cache_data for the DataFrame
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;streamlit&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;

&lt;span class="c1"&gt;# WRONG:
# @st.cache_data           -&amp;gt; tries to pickle a live connection (fails / re-opens)
# def get_connection(): ...
# @st.cache_resource       -&amp;gt; shares ONE DataFrame object; mutations leak across users
# def load_orders(): ...
&lt;/span&gt;
&lt;span class="c1"&gt;# RIGHT:
&lt;/span&gt;&lt;span class="nd"&gt;@st.cache_resource&lt;/span&gt;                      &lt;span class="c1"&gt;# the connection is a shared singleton
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_connection&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;open_warehouse_connection&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="nd"&gt;@st.cache_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ttl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;600&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                 &lt;span class="c1"&gt;# each caller gets a fresh copy of the data
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;load_orders&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;conn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_connection&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_sql&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT * FROM orders WHERE region = &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;symptom&lt;/th&gt;
&lt;th&gt;wrong decorator&lt;/th&gt;
&lt;th&gt;what actually happens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;app is slow&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;get_connection&lt;/code&gt; under &lt;code&gt;cache_data&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;connection is not a picklable value; caching misbehaves, so it re-opens each rerun&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;user B sees user A's data&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;load_orders&lt;/code&gt; under &lt;code&gt;cache_resource&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;the DataFrame is a shared singleton — no copy — so B reads A's filtered frame&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;fixed: fast&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;get_connection&lt;/code&gt; under &lt;code&gt;cache_resource&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;one connection, shared, reused every rerun&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;fixed: isolated&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;load_orders&lt;/code&gt; under &lt;code&gt;cache_data&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;each region's result is a separate copy, keyed on &lt;code&gt;region&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;st.cache_data&lt;/code&gt; is designed for &lt;strong&gt;serializable return values&lt;/strong&gt; and hands back a copy; a live DB connection is neither serializable nor safe to copy, so it belongs in &lt;code&gt;cache_resource&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;st.cache_resource&lt;/code&gt; shares &lt;strong&gt;one object&lt;/strong&gt; across all sessions; a DataFrame stored there is mutated in place by whoever touches it, leaking one user's view into another's.&lt;/li&gt;
&lt;li&gt;Swapping them restores both properties: the connection is opened once and shared, and each query result is an isolated, argument-keyed copy.&lt;/li&gt;
&lt;li&gt;Adding &lt;code&gt;ttl=600&lt;/code&gt; to &lt;code&gt;load_orders&lt;/code&gt; bounds staleness without changing correctness.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;function&lt;/th&gt;
&lt;th&gt;correct decorator&lt;/th&gt;
&lt;th&gt;guarantee&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_connection&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;st.cache_resource&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;one shared, reused connection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;load_orders&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;st.cache_data&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;per-arg copy, no cross-user leak&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;cache_data semantics&lt;/strong&gt;&lt;/strong&gt; — memoizes a function's return value keyed on its arguments and returns a fresh copy each call, which is exactly right for query results and computed frames.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;cache_resource semantics&lt;/strong&gt;&lt;/strong&gt; — caches a single global object with no copy, which is exactly right for connections and models that are expensive to build and meant to be shared.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Copy vs share&lt;/strong&gt;&lt;/strong&gt; — the copy in &lt;code&gt;cache_data&lt;/code&gt; is what isolates users; the shared instance in &lt;code&gt;cache_resource&lt;/code&gt; is what makes a connection reusable — mixing them swaps safety for a leak and speed for a stall.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;TTL invalidation&lt;/strong&gt;&lt;/strong&gt; — a time-to-live bounds staleness so a dashboard refreshes on its own, and &lt;code&gt;.clear()&lt;/code&gt; lets a write path evict immediately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — a cache hit is O(1) plus a hash of the arguments, turning an O(query) round trip on every rerun into one round trip per distinct argument set.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;DataFrames&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — dataframe-basics&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;DataFrame caching and memoization problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/dataframe-basics" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Analysis&lt;/span&gt;
&lt;span&gt;Topic — data-analysis&lt;/span&gt;
&lt;strong&gt;Query-and-aggregate analysis problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Displaying data &amp;amp; connecting to warehouses
&lt;/h2&gt;
&lt;h3&gt;
  
  
  From DataFrame to dashboard is one call — &lt;code&gt;st.dataframe&lt;/code&gt;, the built-in charts, and &lt;code&gt;st.connection&lt;/code&gt; for the warehouse
&lt;/h3&gt;

&lt;p&gt;Once you have a DataFrame, Streamlit gives you display primitives that need no plotting or web code, and a first-class way to &lt;em&gt;get&lt;/em&gt; that DataFrame from a warehouse. The framing an interviewer wants: &lt;strong&gt;you render data with &lt;code&gt;st.dataframe&lt;/code&gt; and the &lt;code&gt;st.*_chart&lt;/code&gt; family, and you source it with &lt;code&gt;st.connection&lt;/code&gt;, whose &lt;code&gt;conn.query(...)&lt;/code&gt; caches results for you&lt;/strong&gt;. No ORM, no cursor management, no manual connection pooling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Display primitives, by intent.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.dataframe(df)&lt;/code&gt; — interactive.&lt;/strong&gt; A sortable, scrollable, searchable grid. Style it with &lt;code&gt;column_config&lt;/code&gt; (format a number as currency, render a URL as a link, show a progress bar in a cell) and enable row selection when you need the user to pick rows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.data_editor(df)&lt;/code&gt; — editable.&lt;/strong&gt; The same grid, but the user can edit cells; it returns the edited DataFrame, which is how you build a lightweight config or mapping editor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.table(df)&lt;/code&gt; — static.&lt;/strong&gt; Renders the whole frame at once, no interactivity — good for small, fixed summaries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.metric(label, value, delta)&lt;/code&gt; — a KPI tile.&lt;/strong&gt; A big number with an optional up/down delta; the building block of the header row on any dashboard.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Built-in charts, zero imports.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.line_chart&lt;/code&gt; / &lt;code&gt;st.area_chart&lt;/code&gt; / &lt;code&gt;st.bar_chart&lt;/code&gt; / &lt;code&gt;st.scatter_chart&lt;/code&gt;.&lt;/strong&gt; Pass a DataFrame plus &lt;code&gt;x=&lt;/code&gt; and &lt;code&gt;y=&lt;/code&gt; (and optional &lt;code&gt;color=&lt;/code&gt; to split series) and Streamlit draws an interactive chart with no Matplotlib or Plotly import. Ideal for latency-over-time and volume-by-pipeline views.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escape hatch when you need it.&lt;/strong&gt; For full control, &lt;code&gt;st.plotly_chart(fig)&lt;/code&gt;, &lt;code&gt;st.altair_chart(chart)&lt;/code&gt;, and &lt;code&gt;st.pyplot(fig)&lt;/code&gt; accept figures from those libraries — you drop down only when the built-ins are not enough.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Connecting to a warehouse with &lt;code&gt;st.connection&lt;/code&gt;.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One call, configured by secrets.&lt;/strong&gt; &lt;code&gt;conn = st.connection("warehouse", type="sql")&lt;/code&gt; reads its URL/credentials from &lt;code&gt;.streamlit/secrets.toml&lt;/code&gt; under &lt;code&gt;[connections.warehouse]&lt;/code&gt;, so no credentials live in code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;conn.query(sql, ttl=...)&lt;/code&gt; caches results.&lt;/strong&gt; The &lt;code&gt;SQLConnection.query&lt;/code&gt; method runs the SQL and memoizes the DataFrame with a built-in TTL — it is &lt;code&gt;st.cache_data&lt;/code&gt; under the hood, so you get caching without decorating anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Typed connections.&lt;/strong&gt; &lt;code&gt;type="sql"&lt;/code&gt; covers any SQLAlchemy URL (Postgres, MySQL, DuckDB, SQLite); &lt;code&gt;type="snowflake"&lt;/code&gt; uses the Snowflake connector; BigQuery is reachable via the SQLAlchemy dialect or a &lt;code&gt;cache_resource&lt;/code&gt; client. Parameterise with &lt;code&gt;params=&lt;/code&gt; — never f-string user input into SQL.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fotqbc8ozmwywcigeudu0.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fotqbc8ozmwywcigeudu0.jpeg" alt="Iconographic Streamlit data-connection diagram — a warehouse cylinder feeding st.connection, a cached conn.query returning a DataFrame, and that DataFrame rendered as an interactive table, a line chart, and a KPI metric." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — query the warehouse and render three ways
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The bread-and-butter dashboard cell queries the warehouse once through &lt;code&gt;st.connection&lt;/code&gt;, then feeds the single result DataFrame into a metric, a chart, and a table. One query, three renders — and the query is cached, so reruns are free until the TTL lapses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Using &lt;code&gt;st.connection&lt;/code&gt;, pull daily loaded-row counts for the last 14 days and show a total-rows metric, a line chart of the trend, and the raw table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt; A &lt;code&gt;daily_loads&lt;/code&gt; table with &lt;code&gt;day&lt;/code&gt; and &lt;code&gt;rows_loaded&lt;/code&gt; columns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;streamlit&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;

&lt;span class="n"&gt;conn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;warehouse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sql&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    SELECT day, rows_loaded
    FROM daily_loads
    WHERE day &amp;gt; current_date - 14
    ORDER BY day
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;ttl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;600&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                            &lt;span class="c1"&gt;# cache the result for 10 minutes
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;metric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Rows loaded (14d)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;rows_loaded&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;line_chart&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;day&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rows_loaded&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dataframe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;use_container_width&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;st.connection("warehouse", type="sql")&lt;/code&gt; builds (or reuses) a connection from the &lt;code&gt;[connections.warehouse]&lt;/code&gt; block in &lt;code&gt;secrets.toml&lt;/code&gt;. &lt;code&gt;conn.query(..., ttl=600)&lt;/code&gt; executes the SQL and caches the DataFrame for ten minutes, so any rerun within that window skips the round trip. The one &lt;code&gt;df&lt;/code&gt; then drives three widgets: &lt;code&gt;st.metric&lt;/code&gt; shows the summed total with a thousands separator, &lt;code&gt;st.line_chart&lt;/code&gt; plots &lt;code&gt;rows_loaded&lt;/code&gt; against &lt;code&gt;day&lt;/code&gt;, and &lt;code&gt;st.dataframe&lt;/code&gt; renders the sortable grid. Nothing here imports a database driver or a charting library directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;widget&lt;/th&gt;
&lt;th&gt;shows&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;st.metric&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"Rows loaded (14d) — 128,400"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;st.line_chart&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;14-point trend line of daily volume&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;st.dataframe&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;sortable 14-row table&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Query once with &lt;code&gt;conn.query(..., ttl=...)&lt;/code&gt; and reuse the one DataFrame for every widget; do not run a separate query per chart, and never build SQL by concatenating user input — pass &lt;code&gt;params=&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Streamlit interview question on warehouse connections
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A dashboard calls &lt;code&gt;pd.read_sql(build_sql(region), psycopg2.connect(DSN))&lt;/code&gt; at the top of the script. It re-opens a connection and re-queries on every slider move, and the DSN is hard-coded. Rewrite it the Streamlit way so the connection is reused, the query is cached per region, and credentials are out of the code.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using st.connection with a parameterised, cached query
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;streamlit&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;

&lt;span class="c1"&gt;# BEFORE: new connection + query on every rerun, credentials in code, SQL injection risk
# conn = psycopg2.connect("host=... user=... password=...")
# df = pd.read_sql(f"SELECT * FROM runs WHERE region = '{region}'", conn)
&lt;/span&gt;
&lt;span class="n"&gt;conn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;warehouse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sql&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# credentials from secrets.toml
&lt;/span&gt;
&lt;span class="n"&gt;region&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;selectbox&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apac&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT pipeline, status, rows_loaded FROM runs WHERE region = :region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;          &lt;span class="c1"&gt;# bound param, not string interpolation
&lt;/span&gt;    &lt;span class="n"&gt;ttl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;600&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                             &lt;span class="c1"&gt;# cache per region for 10 minutes
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dataframe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;use_container_width&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;interaction&lt;/th&gt;
&lt;th&gt;region&lt;/th&gt;
&lt;th&gt;connection&lt;/th&gt;
&lt;th&gt;query executed?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;load&lt;/td&gt;
&lt;td&gt;us&lt;/td&gt;
&lt;td&gt;opened once, cached&lt;/td&gt;
&lt;td&gt;yes (cold)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pick eu&lt;/td&gt;
&lt;td&gt;eu&lt;/td&gt;
&lt;td&gt;reused&lt;/td&gt;
&lt;td&gt;yes (new param)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pick us again&lt;/td&gt;
&lt;td&gt;us&lt;/td&gt;
&lt;td&gt;reused&lt;/td&gt;
&lt;td&gt;no (cache hit)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pick us, 11 min later&lt;/td&gt;
&lt;td&gt;us&lt;/td&gt;
&lt;td&gt;reused&lt;/td&gt;
&lt;td&gt;yes (ttl expired)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;st.connection(...)&lt;/code&gt; builds the connection once and caches it (via &lt;code&gt;cache_resource&lt;/code&gt; internally), so no rerun re-opens it.&lt;/li&gt;
&lt;li&gt;Credentials move to &lt;code&gt;.streamlit/secrets.toml&lt;/code&gt; under &lt;code&gt;[connections.warehouse]&lt;/code&gt;, so the DSN is out of source control.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;conn.query(sql, params={"region": region}, ttl=600)&lt;/code&gt; binds &lt;code&gt;region&lt;/code&gt; as a parameter — closing the SQL-injection hole — and memoizes the result per region for ten minutes.&lt;/li&gt;
&lt;li&gt;Re-selecting a region within the TTL is a cache hit, so the warehouse is queried once per distinct region per ten-minute window instead of once per interaction.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;property&lt;/th&gt;
&lt;th&gt;before&lt;/th&gt;
&lt;th&gt;after&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;connections opened&lt;/td&gt;
&lt;td&gt;one per rerun&lt;/td&gt;
&lt;td&gt;one, reused&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;queries per repeat interaction&lt;/td&gt;
&lt;td&gt;one each time&lt;/td&gt;
&lt;td&gt;one per region per TTL&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;credentials&lt;/td&gt;
&lt;td&gt;hard-coded&lt;/td&gt;
&lt;td&gt;in &lt;code&gt;secrets.toml&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;injection risk&lt;/td&gt;
&lt;td&gt;yes (f-string)&lt;/td&gt;
&lt;td&gt;no (bound param)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;st.connection&lt;/strong&gt;&lt;/strong&gt; — wraps connection construction in a cached factory, so the expensive open happens once and every rerun reuses the same handle.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cached query&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;conn.query(..., ttl=...)&lt;/code&gt; is &lt;code&gt;cache_data&lt;/code&gt; under the hood, keying the DataFrame on the SQL and params, so identical interactions are free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Bound parameters&lt;/strong&gt;&lt;/strong&gt; — passing &lt;code&gt;params=&lt;/code&gt; sends values separately from the SQL text, which both enables correct cache keying and eliminates injection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Secrets separation&lt;/strong&gt;&lt;/strong&gt; — reading credentials from &lt;code&gt;secrets.toml&lt;/code&gt; keeps them out of the repo and lets the same code run against dev and prod by swapping the file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the query cost collapses from O(interactions) round trips to O(distinct params per TTL), the single biggest performance win in a Streamlit data app.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Analysis&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — data-analysis&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Query, aggregate and visualize problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;SLA&lt;/span&gt;
&lt;span&gt;Topic — sla-monitoring&lt;/span&gt;
&lt;strong&gt;Freshness and SLA-metric problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/sla-monitoring" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Building a pipeline-health / SLA dashboard
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Metrics, a freshness table, a backfill button — assembling the primitives into an SLA dashboard with forms and fragments
&lt;/h3&gt;

&lt;p&gt;Everything so far converges here: a real internal tool a data team uses. A pipeline-health dashboard shows &lt;strong&gt;KPI metrics up top, a freshness table with SLA breaches flagged, and controls that &lt;em&gt;do&lt;/em&gt; something — a form that triggers a backfill&lt;/strong&gt;. Two more primitives finish it: &lt;code&gt;st.form&lt;/code&gt; to batch inputs so the script does not rerun on every keystroke, and &lt;code&gt;st.fragment&lt;/code&gt; to refresh part of the page without re-running the whole thing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The layout primitives.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.columns&lt;/code&gt; for the KPI row.&lt;/strong&gt; &lt;code&gt;c1, c2, c3 = st.columns(3)&lt;/code&gt; then &lt;code&gt;c1.metric(...)&lt;/code&gt; places metrics side by side. &lt;code&gt;st.metric("SLA met", "96%", delta="-2%")&lt;/code&gt; shows the value and a coloured delta.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.dataframe&lt;/code&gt; with &lt;code&gt;column_config&lt;/code&gt; for the freshness table.&lt;/strong&gt; Format &lt;code&gt;last_run&lt;/code&gt; as a timestamp, render &lt;code&gt;minutes_late&lt;/code&gt; with a progress bar, and use row styling (via a Styler or a status column) to make breaches obvious.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.sidebar&lt;/code&gt; and &lt;code&gt;st.tabs&lt;/code&gt; for structure.&lt;/strong&gt; Filters go in &lt;code&gt;st.sidebar&lt;/code&gt;; separate an "Overview" tab from a "Backfill" tab with &lt;code&gt;st.tabs([...])&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Forms and buttons — controls that act.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.form&lt;/code&gt; batches input.&lt;/strong&gt; Widgets inside &lt;code&gt;with st.form("backfill"):&lt;/code&gt; do &lt;strong&gt;not&lt;/strong&gt; trigger a rerun individually; the script reruns only when the user clicks &lt;code&gt;st.form_submit_button(...)&lt;/code&gt;. This is essential for a backfill panel with several inputs — you do not want a rerun (or a job) fired on every field change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The submit button gates the action.&lt;/strong&gt; &lt;code&gt;if st.form_submit_button("Run backfill"):&lt;/code&gt; is &lt;code&gt;True&lt;/code&gt; on the submitting rerun, so you kick the job inside that block. Guard destructive actions with a confirmation checkbox or a typed-in pipeline name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Long jobs get status feedback.&lt;/strong&gt; Wrap the work in &lt;code&gt;with st.status("Backfilling...", expanded=True) as s:&lt;/code&gt; and update it, or use &lt;code&gt;st.spinner&lt;/code&gt; / &lt;code&gt;st.progress&lt;/code&gt;, so the user sees progress instead of a frozen page.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Auto-refresh without re-running everything.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;@st.fragment&lt;/code&gt; for partial reruns.&lt;/strong&gt; Decorate a function with &lt;code&gt;@st.fragment&lt;/code&gt; and only that block reruns when its own widgets change — the rest of the page stays put. This keeps a heavy dashboard responsive when one control should not re-query everything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;run_every&lt;/code&gt; for a live tile.&lt;/strong&gt; &lt;code&gt;@st.fragment(run_every="30s")&lt;/code&gt; re-runs that fragment on a timer, so a freshness table or a "last updated" clock refreshes itself without a full-page reload.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;st.rerun()&lt;/code&gt; to force a cycle.&lt;/strong&gt; After a backfill writes new rows, call &lt;code&gt;load.clear()&lt;/code&gt; then &lt;code&gt;st.rerun()&lt;/code&gt; to evict the cache and redraw with fresh data immediately, rather than waiting for the TTL.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Deployment, briefly.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;streamlit run app.py&lt;/code&gt;&lt;/strong&gt; locally; behind a reverse proxy or on &lt;strong&gt;Streamlit Community Cloud&lt;/strong&gt; (point it at a GitHub repo, set secrets in the dashboard) for zero-ops sharing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Docker&lt;/strong&gt; for self-hosting: base image, &lt;code&gt;pip install -r requirements.txt&lt;/code&gt;, &lt;code&gt;EXPOSE 8501&lt;/code&gt;, &lt;code&gt;CMD ["streamlit", "run", "app.py"]&lt;/code&gt;. Put credentials in &lt;code&gt;secrets.toml&lt;/code&gt; or environment variables, never in the image.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3l589mune8x7bh70i01m.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3l589mune8x7bh70i01m.jpeg" alt="Iconographic Streamlit SLA-dashboard diagram — a KPI metric row, a freshness table with a red SLA-breach row highlighted, a backfill form with a submit button, and an auto-refreshing fragment loop." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — a freshness table that flags SLA breaches
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The heart of a pipeline-health dashboard is a freshness table: for each pipeline, how long since its last successful run, and is that within its SLA. The logic is a per-row comparison of lateness against a threshold, rendered so breaches jump out. This is the view an on-call engineer opens first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Given a DataFrame of pipelines with &lt;code&gt;minutes_since_run&lt;/code&gt; and an &lt;code&gt;sla_minutes&lt;/code&gt; threshold, compute a &lt;code&gt;status&lt;/code&gt; of OK or BREACH and show a metric for the breach count plus the flagged table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;pipeline&lt;/th&gt;
&lt;th&gt;minutes_since_run&lt;/th&gt;
&lt;th&gt;sla_minutes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;orders_el&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;events_el&lt;/td&gt;
&lt;td&gt;130&lt;/td&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;dim_refresh&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;45&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;streamlit&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;

&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;load_freshness&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                   &lt;span class="c1"&gt;# cached conn.query(...) from earlier
&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;apply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;BREACH&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;minutes_since_run&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sla_minutes&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OK&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;breaches&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;BREACH&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;

&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;metric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SLA breaches&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;breaches&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dataframe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;style&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;apply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;background-color: #fee2e2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;BREACH&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;use_container_width&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;load_freshness()&lt;/code&gt; returns the cached freshness DataFrame. The &lt;code&gt;apply&lt;/code&gt; compares &lt;code&gt;minutes_since_run&lt;/code&gt; to each pipeline's own &lt;code&gt;sla_minutes&lt;/code&gt; and writes &lt;code&gt;BREACH&lt;/code&gt; or &lt;code&gt;OK&lt;/code&gt; into a new &lt;code&gt;status&lt;/code&gt; column. &lt;code&gt;breaches&lt;/code&gt; counts the flagged rows for the KPI tile. &lt;code&gt;st.metric&lt;/code&gt; surfaces that count up top, and &lt;code&gt;st.dataframe&lt;/code&gt; renders the table with a row-level Styler that tints breach rows red, so the one late pipeline is impossible to miss.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;pipeline&lt;/th&gt;
&lt;th&gt;minutes_since_run&lt;/th&gt;
&lt;th&gt;sla_minutes&lt;/th&gt;
&lt;th&gt;status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;orders_el&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;td&gt;OK&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;events_el&lt;/td&gt;
&lt;td&gt;130&lt;/td&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;td&gt;BREACH&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;dim_refresh&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;45&lt;/td&gt;
&lt;td&gt;OK&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The metric reads &lt;strong&gt;SLA breaches — 1&lt;/strong&gt;, and the &lt;code&gt;events_el&lt;/code&gt; row is highlighted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Compare lateness against a &lt;em&gt;per-pipeline&lt;/em&gt; SLA, not one global threshold; render the breach as both a headline metric and a highlighted row so the summary and the detail agree.&lt;/p&gt;

&lt;h3&gt;
  
  
  Streamlit interview question on triggering a backfill safely
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Product wants a button on the dashboard that backfills a chosen pipeline over a chosen date range. A naive &lt;code&gt;if st.button("Backfill"): run_backfill(...)&lt;/code&gt; outside a form fires on stray reruns and offers no confirmation. Build a safe backfill control that only runs on explicit submit and refreshes the data afterward.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using st.form, a submit button, and a post-run rerun
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;streamlit&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;form&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;backfill&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;pipeline&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;selectbox&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Pipeline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;orders_el&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;events_el&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dim_refresh&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;date_input&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Date range&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt;
    &lt;span class="n"&gt;confirm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;checkbox&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;I understand this reprocesses data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;submitted&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;form_submit_button&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Run backfill&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;submitted&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;confirm&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Backfilling &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;pipeline&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;expanded&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;run_backfill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pipeline&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="c1"&gt;# your job trigger
&lt;/span&gt;        &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;label&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Backfill complete&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;complete&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;load_freshness&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;clear&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                       &lt;span class="c1"&gt;# evict stale cached data
&lt;/span&gt;    &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;rerun&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                                   &lt;span class="c1"&gt;# redraw with fresh numbers
&lt;/span&gt;&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;submitted&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;confirm&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;warning&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Tick the confirmation box before running a backfill.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;user action&lt;/th&gt;
&lt;th&gt;reruns fired&lt;/th&gt;
&lt;th&gt;backfill runs?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;change pipeline dropdown (in form)&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;no — form defers rerun&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pick a date range (in form)&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;no — still deferred&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;click "Run backfill", confirm unticked&lt;/td&gt;
&lt;td&gt;1 (on submit)&lt;/td&gt;
&lt;td&gt;no — warning shown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tick confirm, click "Run backfill"&lt;/td&gt;
&lt;td&gt;1 (on submit)&lt;/td&gt;
&lt;td&gt;yes, then cache cleared + rerun&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Widgets inside &lt;code&gt;with st.form(...)&lt;/code&gt; do not trigger a rerun on change, so choosing a pipeline and a date range fires &lt;strong&gt;no&lt;/strong&gt; premature runs and &lt;strong&gt;no&lt;/strong&gt; stray job.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;st.form_submit_button&lt;/code&gt; reruns the script once and returns &lt;code&gt;True&lt;/code&gt; only on that submitting run, so the job can only start from an explicit click.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;and confirm&lt;/code&gt; gate requires the checkbox, turning an accidental click into a harmless warning instead of a reprocessing job.&lt;/li&gt;
&lt;li&gt;After the job, &lt;code&gt;load_freshness.clear()&lt;/code&gt; evicts the cached freshness data and &lt;code&gt;st.rerun()&lt;/code&gt; forces an immediate redraw, so the dashboard reflects the backfill without waiting for the TTL.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;guarantee&lt;/th&gt;
&lt;th&gt;mechanism&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;no run on field edits&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;st.form&lt;/code&gt; defers reruns to submit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;no run without confirmation&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;and confirm&lt;/code&gt; gate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;fresh data after run&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.clear()&lt;/code&gt; + &lt;code&gt;st.rerun()&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;visible progress&lt;/td&gt;
&lt;td&gt;&lt;code&gt;st.status(...)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;st.form batching&lt;/strong&gt;&lt;/strong&gt; — grouping inputs in a form suppresses the per-widget rerun, so multi-field controls neither thrash the script nor trigger side effects until submit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Submit-gated action&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;st.form_submit_button&lt;/code&gt; is the only trigger, and its &lt;code&gt;True&lt;/code&gt; is confined to the submitting rerun, so a job starts exactly once per intentional click.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Confirmation gate&lt;/strong&gt;&lt;/strong&gt; — requiring a checkbox before a destructive action converts the transient button click into a two-step, mistake-resistant flow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cache clear + rerun&lt;/strong&gt;&lt;/strong&gt; — evicting the memoized query and forcing a rerun makes the freshly written rows appear immediately instead of on the next TTL boundary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the control adds O(1) work and one extra rerun per backfill, a negligible price for making a data-mutating action deliberate and its result immediately visible.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;SLA&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — sla-monitoring&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;SLA-breach and freshness-dashboard problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/sla-monitoring" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Pipelines&lt;/span&gt;
&lt;span&gt;Topic — pipelines&lt;/span&gt;
&lt;strong&gt;Backfill and pipeline-control problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/pipelines" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — Streamlit recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Minimal app.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;streamlit&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;title&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;It renders.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="c1"&gt;# run with: streamlit run app.py
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Widget + session_state counter.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;button&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;＋&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;on_click&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;__setitem__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;session_state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;cache_data query helper.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@st.cache_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ttl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;days&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT * FROM runs WHERE day &amp;gt; current_date - :d&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                      &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;d&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;days&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;st.connection SQL query.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;conn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;warehouse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sql&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# creds in .streamlit/secrets.toml
&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT * FROM daily_loads&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ttl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;600&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dataframe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Form with submit (no rerun until submit).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;form&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;region&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;selectbox&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;go&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;form_submit_button&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Load&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;go&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dataframe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;load_region&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Auto-refreshing fragment.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@st.fragment&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;run_every&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;30s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;freshness_tile&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;metric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Last updated&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;now_str&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;   &lt;span class="c1"&gt;# only this block reruns on the timer
&lt;/span&gt;&lt;span class="nf"&gt;freshness_tile&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Decorator / primitive picker.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cache a DataFrame / query result&lt;/td&gt;
&lt;td&gt;&lt;code&gt;st.cache_data(ttl=...)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache a connection / engine / model&lt;/td&gt;
&lt;td&gt;&lt;code&gt;st.cache_resource&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Keep a value across reruns&lt;/td&gt;
&lt;td&gt;&lt;code&gt;st.session_state&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch several inputs, act on submit&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;st.form&lt;/code&gt; + &lt;code&gt;st.form_submit_button&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Refresh part of the page on a timer&lt;/td&gt;
&lt;td&gt;&lt;code&gt;@st.fragment(run_every=...)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is Streamlit and why do data engineers use it?
&lt;/h3&gt;

&lt;p&gt;Streamlit is an open-source Python framework that turns a plain script into an interactive web app — you write &lt;code&gt;import streamlit as st&lt;/code&gt; and a few &lt;code&gt;st.&lt;/code&gt; calls, run &lt;code&gt;streamlit run app.py&lt;/code&gt;, and get a shareable URL. Data engineers use it because it removes the frontend entirely: a pipeline-health dashboard, an SLA monitor, or a backfill console that would take days in Flask plus React takes an afternoon in one Python file. It is built for internal tools and dashboards, not public consumer products.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does Streamlit's rerun model work?
&lt;/h3&gt;

&lt;p&gt;Streamlit re-executes your entire script from top to bottom every time the user interacts with any widget. Widgets are function calls that return their current value on each run, so you use the returned value inline rather than attaching event handlers. The consequence is that ordinary variables reset on every rerun, which is why persistent values must live in &lt;code&gt;st.session_state&lt;/code&gt; and expensive work must be wrapped in a cache decorator.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between st.cache_data and st.cache_resource?
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;st.cache_data&lt;/code&gt; memoizes a function's return value — DataFrames, query results, anything serializable — keyed on the function's arguments, and it hands each caller a fresh copy so users cannot corrupt each other's data. &lt;code&gt;st.cache_resource&lt;/code&gt; caches a single global object such as a database connection, engine, or ML model and gives every caller the same shared instance, with no copy. Use &lt;code&gt;cache_data&lt;/code&gt; for data you compute or fetch and &lt;code&gt;cache_resource&lt;/code&gt; for things you open once and reuse.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I keep values between reruns in Streamlit?
&lt;/h3&gt;

&lt;p&gt;Store them in &lt;code&gt;st.session_state&lt;/code&gt;, a dictionary-like object scoped to the browser session that survives reruns. Seed a key once with a guard — &lt;code&gt;if "count" not in st.session_state: st.session_state.count = 0&lt;/code&gt; — so the rerun does not reset it, and mutate it inside a widget callback such as &lt;code&gt;on_click&lt;/code&gt;. Binding a widget with &lt;code&gt;key="..."&lt;/code&gt; also mirrors its value into &lt;code&gt;st.session_state&lt;/code&gt; automatically.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does Streamlit connect to a data warehouse?
&lt;/h3&gt;

&lt;p&gt;Use &lt;code&gt;st.connection("name", type="sql")&lt;/code&gt;, which reads credentials from &lt;code&gt;.streamlit/secrets.toml&lt;/code&gt; under &lt;code&gt;[connections.name]&lt;/code&gt; so nothing sensitive lives in code. Call &lt;code&gt;conn.query(sql, params={...}, ttl=600)&lt;/code&gt; to run a parameterised query whose DataFrame result is cached for the TTL — it uses &lt;code&gt;st.cache_data&lt;/code&gt; internally, so repeated interactions do not re-hit the warehouse. &lt;code&gt;type="sql"&lt;/code&gt; covers any SQLAlchemy URL (Postgres, MySQL, DuckDB), and &lt;code&gt;type="snowflake"&lt;/code&gt; targets Snowflake directly.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I build an auto-refreshing pipeline dashboard in Streamlit?
&lt;/h3&gt;

&lt;p&gt;Lay out KPI tiles with &lt;code&gt;st.columns&lt;/code&gt; and &lt;code&gt;st.metric&lt;/code&gt;, render a freshness table with &lt;code&gt;st.dataframe&lt;/code&gt; (highlighting SLA breaches), and put action controls like a backfill trigger inside &lt;code&gt;st.form&lt;/code&gt; so they only fire on submit. For live updates, decorate a section with &lt;code&gt;@st.fragment(run_every="30s")&lt;/code&gt; so just that block re-runs on a timer instead of the whole page, and after a data-mutating action call your cached loader's &lt;code&gt;.clear()&lt;/code&gt; followed by &lt;code&gt;st.rerun()&lt;/code&gt; to redraw with fresh numbers immediately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every Streamlit idea above, from the DataFrame that drives `st.dataframe` and `st.line_chart` to the freshness-and-breach logic behind an SLA table, rests on the data-shaping skills you can drill as graded problems. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you compute per-pipeline SLA status from a freshness table?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/data-analysis" rel="noopener noreferrer"&gt;Practice data-analysis problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/sla-monitoring" rel="noopener noreferrer"&gt;SLA-monitoring drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Factless Fact Tables: Modeling Events, Coverage &amp; Eligibility Without Measures</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Thu, 24 Sep 2026 17:35:55 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/factless-fact-tables-modeling-events-coverage-eligibility-without-measures-2h6h</link>
      <guid>https://dev.to/gowthampotureddi/factless-fact-tables-modeling-events-coverage-eligibility-without-measures-2h6h</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;factless fact table&lt;/code&gt;&lt;/strong&gt; is the one design pattern that breaks the rule every beginner learns first — that a fact table is where the numbers live. A factless fact table has no numbers. It is a fact table made almost entirely of foreign keys pointing at dimensions, with not a single amount, quantity, or price column to sum. And yet it is one of the most useful shapes in a star schema, because the presence of the row is itself the fact. A row exists exactly when a student attended a class, when a product was on promotion, when a user viewed a page — and the metric you want is simply how many such rows there are, or, more subtly, which combinations have no row at all.&lt;/p&gt;

&lt;p&gt;That second question is the one factless fact tables answer better than any other structure: &lt;strong&gt;what did &lt;em&gt;not&lt;/em&gt; happen&lt;/strong&gt;. A normal sales fact only records the sales that occurred, so it can never tell you which promoted products failed to sell — those rows were never written. This guide walks through the four ideas an interviewer will actually probe about factless fact tables — why a measureless fact table is legitimate, the event-tracking family you count with &lt;code&gt;COUNT(*)&lt;/code&gt;, the coverage / eligibility family that records the universe of what &lt;em&gt;could&lt;/em&gt; occur, and the anti-join technique that turns the two into an answer for absence — and pairs each with a Solution-Tail interview answer: code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjzn49hzv2aw7t8xfl4ki.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjzn49hzv2aw7t8xfl4ki.jpeg" alt="PipeCode blog header for factless fact tables — bold white headline 'Factless Fact Tables' with subtitle 'events, coverage, eligibility — no measures' and a stylised star-schema scene where a keys-only fact box connects to dimensions on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/factless-fact" rel="noopener noreferrer"&gt;factless fact practice library →&lt;/a&gt;, rehearse the coverage-and-absence patterns on the &lt;a href="https://pipecode.ai/explore/practice/topic/coverage-fact" rel="noopener noreferrer"&gt;coverage fact practice set →&lt;/a&gt;, and lock in your grain decisions on the &lt;a href="https://pipecode.ai/explore/practice/topic/fact-grain" rel="noopener noreferrer"&gt;fact grain practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why a fact table can have no measures at all&lt;/li&gt;
&lt;li&gt;Event-tracking factless facts — the row is the event&lt;/li&gt;
&lt;li&gt;Coverage &amp;amp; eligibility factless facts — what could happen&lt;/li&gt;
&lt;li&gt;Answering "what did NOT happen" with anti-joins&lt;/li&gt;
&lt;li&gt;Grain, keys &amp;amp; design pitfalls for factless facts&lt;/li&gt;
&lt;li&gt;Cheat sheet — factless fact recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why a fact table can have no measures at all
&lt;/h2&gt;

&lt;h3&gt;
  
  
  A factless fact table is a fact table whose only measure is its own existence — that one idea explains every use of the pattern
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;a factless fact table records that an event or a condition occurred by writing a row of foreign keys, and the measure you want is derived from the rows themselves rather than stored in a column&lt;/strong&gt;. Everything about the pattern follows from that. There is no &lt;code&gt;amount&lt;/code&gt;, no &lt;code&gt;quantity&lt;/code&gt;, no &lt;code&gt;price&lt;/code&gt; — there is a set of dimension keys, maybe a date key, maybe a degenerate id, and the fact is that this particular combination of keys came together at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why "factless" is not a contradiction.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The row is the fact.&lt;/strong&gt; In a sales fact you store &lt;code&gt;amount&lt;/code&gt;; the row plus the number is the fact. In a factless fact the row &lt;em&gt;alone&lt;/em&gt; is the fact — one row means "this happened" or "this was true," and the numeric answer is &lt;code&gt;COUNT(*)&lt;/code&gt; over those rows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The measure is implicit, not absent.&lt;/strong&gt; "Factless" is a slightly misleading name: there is a fact — an occurrence or a condition — it simply is not pre-aggregated into a stored number. The measure lives in the cardinality of the table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It still obeys star-schema rules.&lt;/strong&gt; A factless fact table sits at the centre of a star, joins to conformed dimensions on surrogate keys, and has a clearly declared grain. It is a fact table in every structural sense except the measure column.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The two families you must be able to name.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Event tracking.&lt;/strong&gt; One row every time something happens: a student attends a class, a user views a page, a visitor logs in, a patient is admitted. You answer questions by counting rows and counting distinct dimension members.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Coverage / condition (eligibility).&lt;/strong&gt; One row for every combination that is &lt;em&gt;eligible&lt;/em&gt; or &lt;em&gt;in effect&lt;/em&gt;, whether or not anything happened: which products are on promotion this week, which students are enrolled in which course, which sales reps cover which territory. This records the universe of the possible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The two are complementary.&lt;/strong&gt; Event tables tell you what occurred; coverage tables tell you what &lt;em&gt;could have&lt;/em&gt; occurred. You need both to ask "what was possible but never happened."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you say &lt;strong&gt;"the existence of the row is the measure"&lt;/strong&gt; in the first sentence? — senior signal.&lt;/li&gt;
&lt;li&gt;Can you name &lt;strong&gt;both families&lt;/strong&gt; — event tracking &lt;em&gt;and&lt;/em&gt; coverage — unprompted? — required framing.&lt;/li&gt;
&lt;li&gt;Do you reach for a coverage table specifically to answer &lt;strong&gt;"what did not happen"&lt;/strong&gt;, and explain why the event fact alone cannot? — the whole point.&lt;/li&gt;
&lt;li&gt;Do you resist the temptation to add a &lt;strong&gt;dummy &lt;code&gt;count = 1&lt;/code&gt; column&lt;/strong&gt;, and explain that &lt;code&gt;COUNT(*)&lt;/code&gt; already gives it? — a common anti-pattern they probe.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — the same event, with and without a measure
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The quickest way to feel why a factless design is legitimate is to put a conventional sales fact next to an attendance fact. The sales fact has a measure to sum; the attendance fact has nothing to sum, yet both answer real business questions. The attendance fact answers "how many attendances" with &lt;code&gt;COUNT(*)&lt;/code&gt; — the same answer a stored &lt;code&gt;attendances = 1&lt;/code&gt; column would give, minus the redundant column.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Show a normal sales fact row and a factless attendance fact row side by side, and state how you get a number out of each.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;fact table&lt;/th&gt;
&lt;th&gt;keys&lt;/th&gt;
&lt;th&gt;measure column&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fct_sales&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;product_key, date_key, store_key&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;amount&lt;/code&gt; = 42.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fct_attendance&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;student_key, class_key, date_key&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(none)&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- measured fact: sum the stored measure&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;date_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;revenue&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_sales&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;date_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- factless fact: the metric is the row count itself&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;attendances&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_attendance&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;fct_sales&lt;/code&gt; carries &lt;code&gt;amount&lt;/code&gt;, so the metric is &lt;code&gt;SUM(amount)&lt;/code&gt; — the number was written at load time. &lt;code&gt;fct_attendance&lt;/code&gt; carries no numeric column, so there is nothing to &lt;code&gt;SUM&lt;/code&gt;; instead each row &lt;em&gt;is&lt;/em&gt; one attendance, and &lt;code&gt;COUNT(*)&lt;/code&gt; per &lt;code&gt;class_key&lt;/code&gt; returns how many students attended. Both queries are ordinary star-schema aggregations — the only difference is that the factless query counts rows where the measured query sums a column. Adding an &lt;code&gt;attendances INTEGER DEFAULT 1&lt;/code&gt; column to the factless table would let you write &lt;code&gt;SUM(attendances)&lt;/code&gt;, but it would always equal &lt;code&gt;COUNT(*)&lt;/code&gt;, so it stores nothing new.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;query&lt;/th&gt;
&lt;th&gt;grouped by&lt;/th&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;SUM(amount)&lt;/code&gt; on &lt;code&gt;fct_sales&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;date_key&lt;/td&gt;
&lt;td&gt;revenue per day&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;COUNT(*)&lt;/code&gt; on &lt;code&gt;fct_attendance&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;class_key&lt;/td&gt;
&lt;td&gt;attendances per class&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If the business question is "how many times did X happen," you do not need a measure column — you need a row per occurrence and &lt;code&gt;COUNT(*)&lt;/code&gt;. Reach for a measure only when there is a genuine number to add up.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Event-tracking factless facts — the row is the event
&lt;/h2&gt;

&lt;h3&gt;
  
  
  An event factless fact stores one row per occurrence, joins only to dimensions, and is queried with COUNT(*) and COUNT(DISTINCT)
&lt;/h3&gt;

&lt;p&gt;The event-tracking family is the one most people meet first. The grain is "one row each time the event happens," the columns are the foreign keys that describe &lt;em&gt;who / what / when&lt;/em&gt;, and the answer to almost every question is a count. Attendance is the textbook example: a row is written the moment a student attends a class on a date, and the table never stores a quantity because the quantity is always one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Declaring the grain.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Grain is a single sentence.&lt;/strong&gt; "One row per student per class per day attended." Write that sentence before you write the DDL; every column exists to serve it, and no column may be finer or coarser than it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keys describe the event.&lt;/strong&gt; &lt;code&gt;student_key&lt;/code&gt;, &lt;code&gt;class_key&lt;/code&gt;, &lt;code&gt;date_key&lt;/code&gt; — each a surrogate key into a conformed dimension. Together they answer who attended what and when.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No numeric measure.&lt;/strong&gt; There is deliberately nothing to sum. The event's "value" is that it happened.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The metrics you actually compute.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;COUNT(*)&lt;/code&gt; — how many events.&lt;/strong&gt; Attendances per class, page views per article, logins per day. This is the additive workhorse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;COUNT(DISTINCT key)&lt;/code&gt; — how many distinct members.&lt;/strong&gt; Distinct students who attended (headcount, not visit-count), unique visitors, distinct products viewed. This is &lt;em&gt;non-additive&lt;/em&gt; across some dimensions, which matters when you roll up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ratios across dimensions.&lt;/strong&gt; Attendance rate = distinct attendees ÷ enrolled; both numerator and denominator often come from factless tables (the second from a coverage table, section 3).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Degenerate dimensions on event facts.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The event id has no dimension.&lt;/strong&gt; A &lt;code&gt;session_id&lt;/code&gt;, &lt;code&gt;page_view_id&lt;/code&gt;, or &lt;code&gt;admission_no&lt;/code&gt; that identifies the specific event but has no attributes of its own is stored &lt;em&gt;in the fact table&lt;/em&gt; as a degenerate dimension — a key with no dimension to join to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It enables &lt;code&gt;COUNT(DISTINCT event_id)&lt;/code&gt;.&lt;/strong&gt; When a single logical event can emit multiple rows, the degenerate id lets you count events, not rows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is not a measure.&lt;/strong&gt; A degenerate dimension is still a key; it never gets summed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi75mjc24iliptfgg8ocs.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi75mjc24iliptfgg8ocs.jpeg" alt="Iconographic event-tracking factless fact diagram — one row per occurrence such as a student attending a class, carrying only foreign keys, with COUNT(*) shown as the additive metric and a degenerate event id chip." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — a student-attendance factless fact
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The canonical event factless fact is class attendance. Each time a student shows up to a class on a given day, one row is written with three foreign keys and nothing else. The table can grow to billions of rows and still stores not one number, because every question about it is a count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Model an attendance factless fact and count (a) total attendances per class and (b) distinct students who attended each class.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;student_key&lt;/th&gt;
&lt;th&gt;class_key&lt;/th&gt;
&lt;th&gt;date_key&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;20260902&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;101&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fct_attendance&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;student_key&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_student&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;class_key&lt;/span&gt;   &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_class&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;date_key&lt;/span&gt;    &lt;span class="nb"&gt;INT&lt;/span&gt;    &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;-- no measure column: the row is the attendance&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                     &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;attendances&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;student_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;distinct_students&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_attendance&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The table holds only foreign keys, so a row is the atomic statement "this student attended this class on this date." For &lt;code&gt;class_key = 100&lt;/code&gt; there are three rows (student 1 twice on different days, student 2 once), so &lt;code&gt;COUNT(*) = 3&lt;/code&gt; attendances. But only two &lt;em&gt;distinct&lt;/em&gt; students attended, so &lt;code&gt;COUNT(DISTINCT student_key) = 2&lt;/code&gt;. The two numbers answer different questions — visit-count vs headcount — and both fall out of the same keys-only table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;class_key&lt;/th&gt;
&lt;th&gt;attendances&lt;/th&gt;
&lt;th&gt;distinct_students&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;101&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; For an event fact, keep &lt;code&gt;COUNT(*)&lt;/code&gt; (how many times) and &lt;code&gt;COUNT(DISTINCT member)&lt;/code&gt; (how many members) clearly separated in your head — they diverge exactly when a member can generate more than one event.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on counting events in a factless fact
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You have the &lt;code&gt;fct_attendance&lt;/code&gt; table above and a &lt;code&gt;dim_date&lt;/code&gt; with a &lt;code&gt;week&lt;/code&gt; column. An interviewer asks for weekly attendances &lt;em&gt;and&lt;/em&gt; weekly distinct attendees per class, sorted by class and week. Write the query and explain why the two counts can differ within a week.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using COUNT(*) versus COUNT(DISTINCT) over a joined date dimension
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;week&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                       &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;attendances&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;distinct_attendees&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_attendance&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_date&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;week&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;week&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;student_key&lt;/th&gt;
&lt;th&gt;class_key&lt;/th&gt;
&lt;th&gt;date_key&lt;/th&gt;
&lt;th&gt;week&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;td&gt;2026-W36&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;td&gt;2026-W36&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;20260902&lt;/td&gt;
&lt;td&gt;2026-W36&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;101&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;td&gt;2026-W36&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The join to &lt;code&gt;dim_date&lt;/code&gt; attaches a &lt;code&gt;week&lt;/code&gt; label to each attendance row without changing the grain — it is a lookup, one date key maps to one week.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;GROUP BY week, class_key&lt;/code&gt; buckets the rows; class 100 in week W36 has three rows.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;COUNT(*)&lt;/code&gt; returns 3 attendances for that bucket — student 1 attended on two days, counted twice, which is correct for "visits."&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;COUNT(DISTINCT student_key)&lt;/code&gt; returns 2 — student 1's two visits collapse to one person, which is correct for "unique attendees."&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;week&lt;/th&gt;
&lt;th&gt;class_key&lt;/th&gt;
&lt;th&gt;attendances&lt;/th&gt;
&lt;th&gt;distinct_attendees&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-W36&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-W36&lt;/td&gt;
&lt;td&gt;101&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Row-as-event grain&lt;/strong&gt;&lt;/strong&gt; — because one row means one attendance, &lt;code&gt;COUNT(*)&lt;/code&gt; is a fully additive measure you can roll up across any dimension (day → week → term) without a stored column.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;COUNT(DISTINCT) non-additivity&lt;/strong&gt;&lt;/strong&gt; — distinct attendees cannot be summed across weeks (the same student in two weeks would be double-counted), so it must be recomputed at each grain, not aggregated from a lower level.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Dimension join preserves grain&lt;/strong&gt;&lt;/strong&gt; — joining &lt;code&gt;dim_date&lt;/code&gt; to label the week is a many-to-one lookup, so it neither adds nor drops attendance rows; the counts stay correct.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Two questions, one table&lt;/strong&gt;&lt;/strong&gt; — visit-count and headcount are different business metrics that both come from the same factless table, which is exactly why the pattern is so economical.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;COUNT(*)&lt;/code&gt; is O(rows) on a scan or index; &lt;code&gt;COUNT(DISTINCT)&lt;/code&gt; is O(rows) plus a hash/sort on the distinct key, measurably heavier at scale.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;SQL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — factless-fact&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Factless fact event-counting problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/factless-fact" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Modeling&lt;/span&gt;
&lt;span&gt;Topic — event-log&lt;/span&gt;
&lt;strong&gt;Event-log and occurrence-tracking problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/event-log" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. Coverage &amp;amp; eligibility factless facts — what could happen
&lt;/h2&gt;
&lt;h3&gt;
  
  
  A coverage factless fact enumerates the eligible combinations — the universe of the possible — so you can measure the gap between could and did
&lt;/h3&gt;

&lt;p&gt;The second family is the one interviewers use to separate people who memorised "factless = attendance" from people who actually understand the pattern. A coverage (or condition, or eligibility) factless fact does not record events. It records &lt;strong&gt;which combinations are in effect&lt;/strong&gt;: which products are on promotion in which store this week, which students are enrolled in which course this term, which insurance plan covers which procedure. Nothing has to &lt;em&gt;happen&lt;/em&gt; for a coverage row to exist — the row asserts a condition, not an occurrence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why coverage tables exist at all.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Event facts only record what happened.&lt;/strong&gt; A sales fact has rows only for products that sold. It is structurally incapable of telling you which promoted products sold &lt;em&gt;nothing&lt;/em&gt; — those rows were never written.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Coverage records the denominator.&lt;/strong&gt; To ask "what fraction of promoted products sold," you need the full set of promoted products (the denominator) somewhere. That set is the coverage table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The classic case is promotion coverage.&lt;/strong&gt; For every store / product / promotion / day that a promotion is &lt;em&gt;in effect&lt;/em&gt;, write one coverage row — regardless of whether that product sold. Now the promoted-but-unsold products are answerable.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The grain of a coverage table.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One row per eligible combination per time bucket.&lt;/strong&gt; "One row per product per store per promotion per day the promotion runs." The combination, not an event, defines the row.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It can be large.&lt;/strong&gt; Coverage explodes combinatorially — every eligible product in every covered store for every active day — so it is often kept at a coarser grain (per week, per promotion period) or as date-range rows with &lt;code&gt;valid_from&lt;/code&gt; / &lt;code&gt;valid_to&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Snapshot vs range.&lt;/strong&gt; You can materialise coverage as a daily snapshot (one row per active day, easy to join to a date-grain event fact) or as effective-dated ranges (compact, but you join on &lt;code&gt;BETWEEN&lt;/code&gt;). Snapshots trade storage for simpler joins.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What coverage tables model in the wild.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Eligibility.&lt;/strong&gt; Which customers are eligible for an offer, which employees are eligible for a benefit — the set you will later compare against who redeemed or enrolled.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assignment.&lt;/strong&gt; Which sales rep covers which account, which nurse is rostered to which ward — a many-to-many relationship captured as a factless fact, which doubles as a bridge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Applicability.&lt;/strong&gt; Which tax rule applies to which product, which SLA applies to which ticket class — conditions in force that downstream facts are checked against.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcdboovs2yqdiaeu7ngtj.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcdboovs2yqdiaeu7ngtj.jpeg" alt="Iconographic coverage factless fact diagram — one row per eligible combination such as a product on promotion on a date, recording the universe of what could happen, shown next to a sales event fact for contrast." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — a promotion-coverage factless fact
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The archetypal coverage table is promotion coverage. A retailer runs a promotion; some products are included, in some stores, for some dates. The coverage table gets one row for every product-store-date the promotion is in effect — even a product that never rings up a single sale. That is the whole point: coverage captures the products that &lt;em&gt;could&lt;/em&gt; have sold under the promotion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Model a promotion-coverage factless fact and count how many distinct products were on promotion in each store, independent of whether they sold.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;product_key&lt;/th&gt;
&lt;th&gt;store_key&lt;/th&gt;
&lt;th&gt;promo_key&lt;/th&gt;
&lt;th&gt;date_key&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;501&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;502&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fct_promo_coverage&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;product_key&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_product&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;store_key&lt;/span&gt;   &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_store&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;store_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;promo_key&lt;/span&gt;   &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_promotion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;promo_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;date_key&lt;/span&gt;    &lt;span class="nb"&gt;INT&lt;/span&gt;    &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;-- factless: presence means "this product was on promotion here today"&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;store_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;product_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;products_on_promo&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_promo_coverage&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;store_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Each row asserts that a product was on promotion in a store on a date — a &lt;em&gt;condition&lt;/em&gt;, not a sale. Store 1 has three distinct promoted products (500, 501, 502); store 2 has one (500). &lt;code&gt;COUNT(DISTINCT product_key)&lt;/code&gt; per store answers "how wide was the promotion here," a number the sales fact could never give you because the sales fact only knows about products that actually sold.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;store_key&lt;/th&gt;
&lt;th&gt;products_on_promo&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If a question contains the word "eligible," "on offer," "covered," "assigned," or "enrolled" — anything about what is &lt;em&gt;in effect&lt;/em&gt; rather than what &lt;em&gt;occurred&lt;/em&gt; — you are looking at a coverage factless fact, not an event one.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on modeling eligibility as coverage
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A university needs to answer "how many students are enrolled in each course this term" and, later, "which enrolled students never attended." An interviewer asks how you model enrollment and why it must be a separate table from attendance. Give the coverage table and the enrollment count query.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using an enrollment coverage fact separate from the attendance event fact
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fct_enrollment_coverage&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;student_key&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_student&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;course_key&lt;/span&gt;  &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_course&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;course_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;term_key&lt;/span&gt;    &lt;span class="nb"&gt;INT&lt;/span&gt;    &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_term&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;term_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;-- one row per student enrolled in a course for a term (a condition)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;course_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;student_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;enrolled_students&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_enrollment_coverage&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;course_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;student_key&lt;/th&gt;
&lt;th&gt;course_key&lt;/th&gt;
&lt;th&gt;term_key&lt;/th&gt;
&lt;th&gt;meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;202603&lt;/td&gt;
&lt;td&gt;enrolled (may or may not attend)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;202603&lt;/td&gt;
&lt;td&gt;enrolled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;202603&lt;/td&gt;
&lt;td&gt;enrolled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;101&lt;/td&gt;
&lt;td&gt;202603&lt;/td&gt;
&lt;td&gt;enrolled&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Enrollment is a &lt;em&gt;condition&lt;/em&gt; — the student is registered for the course this term whether or not they ever walk in — so it is coverage, not an event.&lt;/li&gt;
&lt;li&gt;Attendance is a separate event fact (&lt;code&gt;fct_attendance&lt;/code&gt;) written only when a student shows up, so it cannot represent enrolled-but-absent students.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;COUNT(DISTINCT student_key)&lt;/code&gt; per course over the coverage table gives the enrolled headcount — course 100 has three enrolled students.&lt;/li&gt;
&lt;li&gt;Keeping enrollment and attendance in &lt;em&gt;two&lt;/em&gt; tables is what later makes "enrolled but never attended" answerable: it is coverage minus events.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;course_key&lt;/th&gt;
&lt;th&gt;enrolled_students&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;101&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Coverage grain&lt;/strong&gt;&lt;/strong&gt; — one row per enrolled student-course-term captures the eligible universe; the row asserts a standing condition, so it exists independently of any attendance event.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Denominator capture&lt;/strong&gt;&lt;/strong&gt; — enrollment is the denominator for attendance-rate and no-show metrics; without a coverage table there is no place that stores the students who &lt;em&gt;could&lt;/em&gt; have attended.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Separation of concerns&lt;/strong&gt;&lt;/strong&gt; — modeling enrollment and attendance as distinct factless facts keeps each at its own honest grain and prevents you from faking absence with outer joins on a single table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Reusable universe&lt;/strong&gt;&lt;/strong&gt; — the same coverage table feeds many questions (enrolled counts, no-show lists, capacity utilisation), which is why it earns its own place in the schema.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the coverage table is O(eligible combinations), which is bounded and usually far smaller than the event stream, so it is cheap to store and index.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Modeling&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — coverage-fact&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Coverage and eligibility modeling problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/coverage-fact" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;SQL&lt;/span&gt;
&lt;span&gt;Topic — factless-fact&lt;/span&gt;
&lt;strong&gt;Factless fact table design problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/factless-fact" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Answering "what did NOT happen" with anti-joins
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Coverage LEFT JOIN events, keep the NULLs — the anti-join is how a factless design turns absence into a queryable result
&lt;/h3&gt;

&lt;p&gt;Here is the payoff that makes the whole pattern worth learning. Business people constantly ask negative questions: which promoted products sold &lt;em&gt;nothing&lt;/em&gt;, which enrolled students &lt;em&gt;never&lt;/em&gt; attended, which eligible customers &lt;em&gt;never&lt;/em&gt; redeemed. An event fact alone cannot answer any of them, because absence leaves no row. The technique is always the same shape: &lt;strong&gt;take the coverage table (what was possible), left-join it to the event table (what happened), and keep only the rows where the event side is NULL&lt;/strong&gt; — those are the non-events.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The anti-join, three equivalent ways.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;LEFT JOIN ... WHERE right.key IS NULL&lt;/code&gt;.&lt;/strong&gt; Join coverage to events on the shared keys; unmatched coverage rows get NULLs on the event side; filter to &lt;code&gt;IS NULL&lt;/code&gt;. The most portable and most readable form.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;NOT EXISTS (correlated subquery)&lt;/code&gt;.&lt;/strong&gt; For each coverage row, check that no matching event row exists. Often the planner's favourite — it can short-circuit on the first match and handles NULLs cleanly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;EXCEPT&lt;/code&gt; / &lt;code&gt;MINUS&lt;/code&gt;.&lt;/strong&gt; Set-difference the coverage key-set minus the event key-set. Elegant when the two sides share exactly the same key columns and you only want the keys.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The correctness traps.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;NOT IN&lt;/code&gt; with NULLs is a landmine.&lt;/strong&gt; &lt;code&gt;WHERE key NOT IN (SELECT key FROM events)&lt;/code&gt; returns &lt;em&gt;no rows at all&lt;/em&gt; if the subquery yields a single NULL, because &lt;code&gt;NOT IN&lt;/code&gt; is three-valued. Prefer &lt;code&gt;NOT EXISTS&lt;/code&gt; or &lt;code&gt;LEFT JOIN ... IS NULL&lt;/code&gt;, which are NULL-safe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Join on the &lt;em&gt;complete&lt;/em&gt; grain.&lt;/strong&gt; The anti-join must match on every key that defines "the same thing." Match promoted products to sales on &lt;code&gt;product_key AND store_key AND date_key&lt;/code&gt;; drop one and you either miss non-events or invent false ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Align the time grain.&lt;/strong&gt; If coverage is per-day and events are per-transaction, aggregate or key both to the same date grain before the anti-join, or a same-day sale in a different hour can be mismatched.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why you need both tables.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Coverage = the universe.&lt;/strong&gt; Without a coverage fact there is no set of "everything that could have happened" to subtract from.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Events = what occurred.&lt;/strong&gt; Without the event fact there is nothing to subtract.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Absence = coverage − events.&lt;/strong&gt; The non-event is precisely the coverage rows with no matching event; that is a set difference you cannot compute from either table alone.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fleumxm0ko8gkrb7a3cuq.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fleumxm0ko8gkrb7a3cuq.jpeg" alt="Iconographic anti-join diagram — a coverage fact LEFT JOINed to an event fact, rows A and C matching, rows B and D finding no match and returning NULL, then surfaced as the non-events such as promoted products that never sold." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — promoted products that never sold
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The definitive "what did not happen" query pairs promotion coverage with a sales event fact. Coverage lists every product on promotion in a store on a day; sales lists every product that actually rang up. Anti-joining coverage to sales on the full key surfaces exactly the promoted products that sold nothing — the rows a plain sales report can never show.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Given &lt;code&gt;fct_promo_coverage&lt;/code&gt; (product/store/date on promotion) and &lt;code&gt;fct_sales&lt;/code&gt; (product/store/date that sold), list the promoted products that had zero sales.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;coverage (product, store, date)&lt;/th&gt;
&lt;th&gt;sales (product, store, date)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;500, 1, 0901&lt;/td&gt;
&lt;td&gt;500, 1, 0901&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;501, 1, 0901&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(no row)&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;502, 1, 0901&lt;/td&gt;
&lt;td&gt;502, 1, 0901&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;500, 2, 0901&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(no row)&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_promo_coverage&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
&lt;span class="k"&gt;LEFT&lt;/span&gt;   &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;fct_sales&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
       &lt;span class="k"&gt;ON&lt;/span&gt;  &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt;
       &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store_key&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store_key&lt;/span&gt;
       &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;    &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;-- promoted but never sold&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The &lt;code&gt;LEFT JOIN&lt;/code&gt; keeps every coverage row and attaches matching sales rows where they exist. Product 500 in store 1 sold, so it matches and is discarded by the &lt;code&gt;IS NULL&lt;/code&gt; filter; product 502 in store 1 likewise matches and drops. Product 501 in store 1 and product 500 in store 2 have no sale on that date, so their sales columns are NULL, and the &lt;code&gt;WHERE s.product_key IS NULL&lt;/code&gt; filter keeps exactly those two — the promoted-but-unsold products.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;product_key&lt;/th&gt;
&lt;th&gt;store_key&lt;/th&gt;
&lt;th&gt;date_key&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;501&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; "What did not happen" is always &lt;code&gt;coverage LEFT JOIN events ... WHERE events.key IS NULL&lt;/code&gt;. If you find yourself trying to answer it from the event table alone, you are missing a coverage fact.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on absence and the NOT IN pitfall
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A colleague wrote &lt;code&gt;WHERE student_key NOT IN (SELECT student_key FROM fct_attendance)&lt;/code&gt; to find enrolled students who never attended, and it returns zero rows even though many students were absent. Explain the bug and give a correct, NULL-safe query for "enrolled but never attended," per course.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using NOT EXISTS instead of NOT IN
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;course_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_enrollment_coverage&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
           &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
           &lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_attendance&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;
           &lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt;
             &lt;span class="k"&gt;AND&lt;/span&gt;  &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;course_key&lt;/span&gt;   &lt;span class="c1"&gt;-- align to the same grain&lt;/span&gt;
       &lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;enrolled (student, course)&lt;/th&gt;
&lt;th&gt;attended?&lt;/th&gt;
&lt;th&gt;NOT EXISTS result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1, 100&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;excluded&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2, 100&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;kept (never attended)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3, 100&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;kept&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3, 101&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;excluded&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The bug: &lt;code&gt;fct_attendance.student_key&lt;/code&gt; contains a NULL (or the subquery can), so &lt;code&gt;NOT IN&lt;/code&gt; evaluates to &lt;code&gt;UNKNOWN&lt;/code&gt; for every row and the whole result collapses to zero — three-valued logic, not a data problem.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;NOT EXISTS&lt;/code&gt; asks a yes/no question per enrolled row — "is there any attendance for this student in this course?" — and NULLs inside the subquery simply fail to match, so they do not poison the result.&lt;/li&gt;
&lt;li&gt;The correlation matches on &lt;em&gt;both&lt;/em&gt; keys (&lt;code&gt;student_key&lt;/code&gt; and course/class), so it checks attendance for the right course, not attendance anywhere.&lt;/li&gt;
&lt;li&gt;Students 2 and 3-in-course-100 have no matching attendance row, so &lt;code&gt;NOT EXISTS&lt;/code&gt; is true and they are correctly returned as never-attended.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;course_key&lt;/th&gt;
&lt;th&gt;student_key&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Anti-join semantics&lt;/strong&gt;&lt;/strong&gt; — "never attended" is a set difference (enrolled minus attended); &lt;code&gt;NOT EXISTS&lt;/code&gt; expresses that difference directly and returns the coverage rows with no event match.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;NULL-safety&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;NOT EXISTS&lt;/code&gt; uses two-valued existence checks, so a NULL in the attendance keys cannot wipe the result the way &lt;code&gt;NOT IN&lt;/code&gt;'s three-valued logic does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Full-grain correlation&lt;/strong&gt;&lt;/strong&gt; — matching on student &lt;em&gt;and&lt;/em&gt; course guarantees you are testing attendance for the same course the student is enrolled in, so you neither miss absentees nor flag the wrong ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Coverage as the driving table&lt;/strong&gt;&lt;/strong&gt; — the query iterates the enrollment universe and probes events, which is the only order that can surface absence; driving from events could never see a student who generated no row.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;NOT EXISTS&lt;/code&gt; is typically an anti-semi-join, O(coverage rows) probing an index on the event keys; far cheaper and safer than materialising a &lt;code&gt;NOT IN&lt;/code&gt; list.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;SQL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — factless-fact&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Anti-join and "what did not happen" problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/factless-fact" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Modeling&lt;/span&gt;
&lt;span&gt;Topic — coverage-fact&lt;/span&gt;
&lt;strong&gt;Coverage-vs-event absence problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/coverage-fact" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Grain, keys &amp;amp; design pitfalls for factless facts
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Declare the grain, count the rows, and never invent a dummy measure — the discipline that keeps a keys-only table honest
&lt;/h3&gt;

&lt;p&gt;A factless fact table is simple to draw and easy to get subtly wrong. Because it has no measure to anchor it, the grain is the only thing keeping the table meaningful, and a sloppy join can silently multiply your "counts." This section is the checklist an interviewer expects you to run through when they say "design me a factless fact."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Grain first, always.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One sentence, no exceptions.&lt;/strong&gt; State the grain — "one row per student per class per day attended" — before any DDL. Every foreign key is there to satisfy that sentence; if a column does not help identify one row, it does not belong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mixed grains corrupt counts.&lt;/strong&gt; If some rows are per-day and others per-week in the same table, &lt;code&gt;COUNT(*)&lt;/code&gt; means nothing. Keep one grain per table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Choose event &lt;em&gt;or&lt;/em&gt; condition, not both.&lt;/strong&gt; A single factless table is either event-tracking or coverage. Cramming both into one table (with a &lt;code&gt;type&lt;/code&gt; flag) destroys the clean &lt;code&gt;COUNT(*)&lt;/code&gt; and forces filters everywhere.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Counting safely.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;COUNT(*)&lt;/code&gt; is additive; &lt;code&gt;COUNT(DISTINCT)&lt;/code&gt; is not.&lt;/strong&gt; You can sum &lt;code&gt;COUNT(*)&lt;/code&gt; up a hierarchy; you cannot sum distinct counts across buckets. Recompute distincts at each grain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Beware join fan-out.&lt;/strong&gt; Joining a factless fact to a dimension that is &lt;em&gt;not&lt;/em&gt; many-to-one (e.g. a bridge, or a mis-keyed dimension with duplicates) multiplies rows and inflates &lt;code&gt;COUNT(*)&lt;/code&gt;. Verify every fact-to-dimension join is many-to-one, or count a degenerate id with &lt;code&gt;COUNT(DISTINCT)&lt;/code&gt; to stay safe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Degenerate ids protect event counts.&lt;/strong&gt; When one logical event can produce several rows, store the event id in the fact and &lt;code&gt;COUNT(DISTINCT event_id)&lt;/code&gt; so a downstream join cannot inflate the metric.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Anti-patterns to name and reject.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The constant &lt;code&gt;1&lt;/code&gt; measure.&lt;/strong&gt; Adding &lt;code&gt;attendances INTEGER DEFAULT 1&lt;/code&gt; so you can "&lt;code&gt;SUM&lt;/code&gt; a fact" is redundant — &lt;code&gt;COUNT(*)&lt;/code&gt; already returns it. It wastes space and invites bugs when someone loads a &lt;code&gt;2&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storing derived attributes on the fact.&lt;/strong&gt; Do not denormalise dimension attributes (student name, class title) into the factless fact; keep it keys-only and join to conformed dimensions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Using a coverage table as an event table (or vice versa).&lt;/strong&gt; Counting coverage rows as if they were events overstates activity; both must exist and be used for what they model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F55frm18bgjmbkhnqivu0.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F55frm18bgjmbkhnqivu0.jpeg" alt="Iconographic factless fact design diagram — a declared-grain lane, a COUNT(*) additivity lane, and a pitfalls lane warning about join fan-out double counting and the constant-one-as-measure anti-pattern." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — how a fan-out join inflates a factless count
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The most common factless-fact bug is a count that comes back too high because a join multiplied the rows. Here an attendance fact is joined to a class-instructor bridge where a class has two co-instructors; the join doubles every attendance row, and a naive &lt;code&gt;COUNT(*)&lt;/code&gt; reports twice the real attendance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; An attendance fact is joined to &lt;code&gt;bridge_class_instructor&lt;/code&gt; (a class can have several instructors) to attribute attendances to instructors. Show how &lt;code&gt;COUNT(*)&lt;/code&gt; inflates and how to count correctly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;attendance (student, class)&lt;/th&gt;
&lt;th&gt;class → instructors&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1, 100&lt;/td&gt;
&lt;td&gt;100 → {A, B}&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2, 100&lt;/td&gt;
&lt;td&gt;100 → {A, B}&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- WRONG: the bridge fans out each attendance into one row per instructor&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;attendances&lt;/span&gt;          &lt;span class="c1"&gt;-- returns 4, not 2&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_attendance&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;bridge_class_instructor&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- RIGHT: count distinct attendance events, immune to the fan-out&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;attendance_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;attendances&lt;/span&gt;   &lt;span class="c1"&gt;-- returns 2&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_attendance&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;bridge_class_instructor&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The bridge maps class 100 to two instructors, so the join turns each of the two attendance rows into two rows (one per instructor) — four rows total. Plain &lt;code&gt;COUNT(*)&lt;/code&gt; now reports 4 attendances, which is wrong: only two students attended. Counting &lt;code&gt;DISTINCT attendance_id&lt;/code&gt; — a degenerate dimension carried on the fact — collapses the fan-out back to the two real events, giving the correct 2 regardless of how many instructors the bridge attaches.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;query&lt;/th&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;th&gt;correct?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;COUNT(*)&lt;/code&gt; after bridge join&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;no (inflated)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;COUNT(DISTINCT attendance_id)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Any join to a non-many-to-one table can inflate &lt;code&gt;COUNT(*)&lt;/code&gt;; carry a degenerate event id on the factless fact and count it distinctly whenever a bridge or multi-valued dimension is in the join path.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on grain and additivity
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You are asked to report "distinct products on promotion per store per week" from &lt;code&gt;fct_promo_coverage&lt;/code&gt; (grain: product/store/promo/day), and a teammate proposes summing a daily &lt;code&gt;products_on_promo&lt;/code&gt; count up to the week. Explain why that is wrong and write the correct weekly query.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using COUNT(DISTINCT) at the target grain instead of summing lower-grain counts
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;week&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;products_on_promo&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_promo_coverage&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_date&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;week&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;product_key&lt;/th&gt;
&lt;th&gt;store_key&lt;/th&gt;
&lt;th&gt;date_key&lt;/th&gt;
&lt;th&gt;week&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;20260901&lt;/td&gt;
&lt;td&gt;2026-W36&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;20260902&lt;/td&gt;
&lt;td&gt;2026-W36&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;501&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;20260902&lt;/td&gt;
&lt;td&gt;2026-W36&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The teammate's plan computes a daily distinct-product count (day 0901 = 1, day 0902 = 2) and sums them to 3 for the week — but product 500 was on promotion on both days and gets counted twice.&lt;/li&gt;
&lt;li&gt;Distinct counts are &lt;strong&gt;non-additive&lt;/strong&gt;: you cannot sum a distinct measure from a lower grain to a higher one without double-counting members that span buckets.&lt;/li&gt;
&lt;li&gt;The correct query recomputes &lt;code&gt;COUNT(DISTINCT product_key)&lt;/code&gt; directly at the week grain, so product 500 (present on two days) is counted once.&lt;/li&gt;
&lt;li&gt;The result is 2 distinct products on promotion in store 1 for the week, not the inflated 3.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;week&lt;/th&gt;
&lt;th&gt;store_key&lt;/th&gt;
&lt;th&gt;products_on_promo&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-W36&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Grain discipline&lt;/strong&gt;&lt;/strong&gt; — recomputing the distinct at the reporting grain respects that the coverage grain is per-day; you never sum across the boundary that would double-count a member.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Non-additive measures&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;COUNT(DISTINCT)&lt;/code&gt; cannot be rolled up like &lt;code&gt;SUM&lt;/code&gt;, so the only correct way to get a weekly distinct is to aggregate the base rows to the week directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Dimension-driven rollup&lt;/strong&gt;&lt;/strong&gt; — joining &lt;code&gt;dim_date&lt;/code&gt; to attach the week lets one query serve any time grain by changing the &lt;code&gt;GROUP BY&lt;/code&gt;, without pre-aggregated intermediate tables that would bake in the additivity bug.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Keys-only integrity&lt;/strong&gt;&lt;/strong&gt; — because the coverage fact stores only keys, there is no stored count to be tempted into summing; the metric is always derived at query time at the right grain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — a single &lt;code&gt;COUNT(DISTINCT)&lt;/code&gt; over the base rows is O(rows) with a hash aggregate, and it avoids the two-pass "aggregate then re-aggregate" that produces the wrong answer anyway.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Grain&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — fact-grain&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Fact-table grain and additivity problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/fact-grain" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;SQL&lt;/span&gt;
&lt;span&gt;Topic — factless-fact&lt;/span&gt;
&lt;strong&gt;Factless fact modeling and pitfall problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/factless-fact" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — factless fact recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Event factless fact (keys only, count the rows).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fct_attendance&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;student_key&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;class_key&lt;/span&gt;   &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;date_key&lt;/span&gt;    &lt;span class="nb"&gt;INT&lt;/span&gt;    &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;attendances&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_attendance&lt;/span&gt; &lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;class_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Coverage / eligibility factless fact (the universe of the possible).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fct_promo_coverage&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;product_key&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;store_key&lt;/span&gt;   &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;promo_key&lt;/span&gt;   &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;date_key&lt;/span&gt;    &lt;span class="nb"&gt;INT&lt;/span&gt;    &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;"What did NOT happen" — LEFT JOIN … IS NULL.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_promo_coverage&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
&lt;span class="k"&gt;LEFT&lt;/span&gt;   &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;fct_sales&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
       &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt;
      &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store_key&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store_key&lt;/span&gt;
      &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;    &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Absence — the NULL-safe NOT EXISTS form.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;course_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_enrollment_coverage&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;fct_attendance&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;
    &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt;
      &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;course_key&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Attendance rate = distinct attendees ÷ enrolled (event ÷ coverage).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;en&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;course_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="k"&gt;at&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt;&lt;span class="p"&gt;)::&lt;/span&gt;&lt;span class="nb"&gt;numeric&lt;/span&gt;
         &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;en&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;attendance_rate&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fct_enrollment_coverage&lt;/span&gt; &lt;span class="n"&gt;en&lt;/span&gt;
&lt;span class="k"&gt;LEFT&lt;/span&gt;   &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;fct_attendance&lt;/span&gt; &lt;span class="k"&gt;at&lt;/span&gt;
       &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;at&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;en&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;student_key&lt;/span&gt;
      &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="k"&gt;at&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;class_key&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;en&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;course_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;en&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;course_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Coverage vs event — which family to build.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question shape&lt;/th&gt;
&lt;th&gt;Family&lt;/th&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"How many times did X happen?"&lt;/td&gt;
&lt;td&gt;event tracking&lt;/td&gt;
&lt;td&gt;&lt;code&gt;COUNT(*)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"How many distinct members did X?"&lt;/td&gt;
&lt;td&gt;event tracking&lt;/td&gt;
&lt;td&gt;&lt;code&gt;COUNT(DISTINCT member)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"What was eligible / on offer / assigned?"&lt;/td&gt;
&lt;td&gt;coverage&lt;/td&gt;
&lt;td&gt;&lt;code&gt;COUNT(DISTINCT combo)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"What was possible but never happened?"&lt;/td&gt;
&lt;td&gt;coverage anti-join event&lt;/td&gt;
&lt;td&gt;&lt;code&gt;LEFT JOIN … IS NULL&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is a factless fact table?
&lt;/h3&gt;

&lt;p&gt;A factless fact table is a fact table that contains only foreign keys to dimensions (plus optionally a date key and a degenerate id) and &lt;strong&gt;no numeric measure column&lt;/strong&gt;. The presence of a row &lt;em&gt;is&lt;/em&gt; the fact: it records that an event occurred or that a condition was in effect. You derive metrics by counting rows with &lt;code&gt;COUNT(*)&lt;/code&gt; and &lt;code&gt;COUNT(DISTINCT ...)&lt;/code&gt; rather than summing a stored measure, so "factless" means "no stored measure," not "no meaning."&lt;/p&gt;

&lt;h3&gt;
  
  
  What are the two types of factless fact tables?
&lt;/h3&gt;

&lt;p&gt;There are two families. &lt;strong&gt;Event tracking&lt;/strong&gt; stores one row every time something happens — a student attends a class, a user views a page, a visitor logs in — and you answer questions by counting those rows. &lt;strong&gt;Coverage (or condition / eligibility)&lt;/strong&gt; stores one row for every combination that is in effect — a product on promotion, a student enrolled in a course, a rep assigned to a territory — regardless of whether anything happened. Event tables tell you what occurred; coverage tables tell you what could have occurred.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you get a measure from a factless fact table?
&lt;/h3&gt;

&lt;p&gt;You count. &lt;code&gt;COUNT(*)&lt;/code&gt; over the rows gives the number of events (attendances, views, logins), and &lt;code&gt;COUNT(DISTINCT some_key)&lt;/code&gt; gives the number of distinct members (unique attendees, unique visitors). Because the grain is "one row per occurrence," the row count &lt;em&gt;is&lt;/em&gt; the measure — adding a constant &lt;code&gt;1&lt;/code&gt; column so you can &lt;code&gt;SUM&lt;/code&gt; it is redundant, since it always equals &lt;code&gt;COUNT(*)&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you answer "what did not happen" in a star schema?
&lt;/h3&gt;

&lt;p&gt;You need two tables: a coverage factless fact for what was possible and an event fact for what occurred. Then you anti-join them — &lt;code&gt;coverage LEFT JOIN events ON &amp;lt;full key&amp;gt; WHERE events.key IS NULL&lt;/code&gt;, or the NULL-safe &lt;code&gt;NOT EXISTS&lt;/code&gt; form. The coverage rows with no matching event are exactly the things that did not happen: promoted products that sold nothing, enrolled students who never attended. An event fact alone can never answer this because absence leaves no row.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is a coverage fact table?
&lt;/h3&gt;

&lt;p&gt;A coverage fact table is a factless fact that enumerates the universe of eligible or in-effect combinations — every product on promotion in every store on every active day, for example — whether or not an event occurred against them. It captures the &lt;em&gt;denominator&lt;/em&gt; you need for rates (attendance rate, redemption rate) and the reference set you anti-join against to find non-events. It typically lives at a per-day snapshot grain or as effective-dated &lt;code&gt;valid_from&lt;/code&gt; / &lt;code&gt;valid_to&lt;/code&gt; ranges.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can a factless fact table have a degenerate dimension?
&lt;/h3&gt;

&lt;p&gt;Yes, and it often should. A degenerate dimension is a key that identifies the specific event — an &lt;code&gt;attendance_id&lt;/code&gt;, &lt;code&gt;session_id&lt;/code&gt;, or &lt;code&gt;admission_no&lt;/code&gt; — but has no attributes of its own, so it is stored in the fact table with no dimension to join to. It lets you &lt;code&gt;COUNT(DISTINCT event_id)&lt;/code&gt; to get an accurate event count even when a join (to a bridge or multi-valued dimension) fans a single event out into several rows. It is still a key, never a measure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every factless-fact idea above, from the event-counting attendance table to the coverage universe and the anti-join for absence, maps to a hands-on practice room where you build the model and the query against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you find the products that were on promotion but never sold?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/factless-fact" rel="noopener noreferrer"&gt;Practice factless-fact problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/coverage-fact" rel="noopener noreferrer"&gt;Coverage-fact drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Bridge Tables &amp; Many-to-Many Dimensions: Modeling Hierarchies and Multi-Valued Attributes</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Wed, 23 Sep 2026 13:52:42 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/bridge-tables-many-to-many-dimensions-modeling-hierarchies-and-multi-valued-attributes-27op</link>
      <guid>https://dev.to/gowthampotureddi/bridge-tables-many-to-many-dimensions-modeling-hierarchies-and-multi-valued-attributes-27op</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;bridge table&lt;/code&gt;&lt;/strong&gt; is the piece of a star schema you reach for the moment a relationship refuses to be many-to-one. A textbook fact row joins to exactly one row in each dimension — one order, one customer, one date. But reality is full of relationships that are many-to-many: a bank account can be owned by several customers, a customer can hold several accounts; a hospital visit can carry several diagnoses; a product can wear several category tags; an employee can sit at any depth beneath a manager in an org chart. A bridge table is the associative table you slot between the two sides so the model can express "many on both ends" without either overwriting data or silently multiplying your numbers.&lt;/p&gt;

&lt;p&gt;The reason interviewers keep asking about bridges is that they sit exactly where correctness breaks. Join a fact to a dimension through a bridge naively and a joint account's balance shows up once per owner, so the bank's total balance inflates — the classic double-counting fan trap. This guide walks the four ideas an interviewer will actually probe — the fact-to-dimension many-to-many and where the bridge goes, the weighting or allocation factor that makes shared measures sum correctly, multi-valued attributes and multi-valued dimensions, and hierarchy bridges / closure tables for variable-depth trees like org charts and bills of materials — and pairs each with a Solution-Tail interview answer: SQL, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxgixia5hea909kqo9ykk.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxgixia5hea909kqo9ykk.jpeg" alt="PipeCode blog header for bridge tables and many-to-many dimensions — bold white headline 'Bridge Tables' with subtitle 'Many-to-Many, Weighting, Hierarchies' and a stylised fact-bridge-dimension scene on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/bridge-table" rel="noopener noreferrer"&gt;bridge-table practice library →&lt;/a&gt;, rehearse the join shapes on the &lt;a href="https://pipecode.ai/explore/practice/topic/many-to-many" rel="noopener noreferrer"&gt;many-to-many practice set →&lt;/a&gt;, and harden your tree queries on the &lt;a href="https://pipecode.ai/explore/practice/topic/recursive-hierarchy" rel="noopener noreferrer"&gt;recursive-hierarchy practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why bridge tables exist — the many-to-many problem&lt;/li&gt;
&lt;li&gt;Bridge tables between a fact and a dimension&lt;/li&gt;
&lt;li&gt;Weighting factors — allocation without double-counting&lt;/li&gt;
&lt;li&gt;Multi-valued attributes &amp;amp; multi-valued dimensions&lt;/li&gt;
&lt;li&gt;Hierarchy bridges &amp;amp; closure tables for ragged trees&lt;/li&gt;
&lt;li&gt;Cheat sheet — bridge &amp;amp; hierarchy recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why bridge tables exist — the many-to-many problem
&lt;/h2&gt;

&lt;h3&gt;
  
  
  A star schema assumes many-to-one — a bridge is what you add when both sides are "many"
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;a bridge table is an associative table placed between two entities that have a many-to-many relationship, holding one row for each valid pair of keys&lt;/strong&gt;. Everything about bridges follows from the fact that a normal dimensional join cannot store "many on both ends." A fact table carries a foreign key to each dimension; that foreign key holds exactly one value, so the fact-to-dimension grain is many facts to one dimension member. The instant a single fact (or a single dimension member) legitimately relates to several members on the other side, that single-valued foreign key runs out of room.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where many-to-many shows up in real models.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Shared ownership.&lt;/strong&gt; One bank account, several account holders (a joint account); one insurance policy, several covered members; one property, several owners.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-valued classification.&lt;/strong&gt; One patient visit, several diagnoses; one support ticket, several tags; one product, several categories or attributes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ragged hierarchies.&lt;/strong&gt; One employee reports through a chain of managers of unknown length; one assembly contains sub-assemblies to arbitrary depth (a bill of materials).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The symmetry that hurts.&lt;/strong&gt; In every case the relationship is many-to-many both ways — a customer also holds many accounts — so neither side can host the foreign key.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The three wrong answers a bridge replaces.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Overwrite (lose data).&lt;/strong&gt; Put a single &lt;code&gt;customer_key&lt;/code&gt; on the account and you can record only one owner; the second holder vanishes. Correct total, wrong coverage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grain explosion (duplicate facts).&lt;/strong&gt; Duplicate the fact once per owner and you have changed the grain of the fact table; every downstream &lt;code&gt;SUM&lt;/code&gt; now double-counts unless every query remembers to guard against it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Comma-list / repeating groups (unqueryable).&lt;/strong&gt; Stuff &lt;code&gt;"C1,C2"&lt;/code&gt; into one column, or add &lt;code&gt;owner_1&lt;/code&gt;, &lt;code&gt;owner_2&lt;/code&gt;, &lt;code&gt;owner_3&lt;/code&gt; columns. Both are anti-patterns: you cannot join, filter, or aggregate on them, and the fixed columns cap the cardinality arbitrarily.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What a bridge actually is.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;An associative table at the grain of the relationship.&lt;/strong&gt; One row per related pair — &lt;code&gt;(account_key, customer_key)&lt;/code&gt; — and nothing else except optional attributes of the &lt;em&gt;relationship&lt;/em&gt; (a weighting factor, an effective-date range, a role like "primary owner").&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A join hub, not a fact.&lt;/strong&gt; A bridge usually has no additive measures of its own; it exists to be joined &lt;em&gt;through&lt;/em&gt;. Its own grain is "one row per pair," which is exactly what lets you fan a query out and — with care — fan it back in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The place correctness is won or lost.&lt;/strong&gt; Because a query passes &lt;em&gt;through&lt;/em&gt; the bridge, the bridge is where you attach the weighting factor that prevents double-counting and the flags that let you pick a rollup level.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do you name the relationship as &lt;strong&gt;many-to-many both ways&lt;/strong&gt; before proposing a table? — the framing that earns trust.&lt;/li&gt;
&lt;li&gt;Do you put the bridge &lt;strong&gt;at the grain of the relationship&lt;/strong&gt; ("one row per account-customer pair"), not at the grain of a fact? — the core idea.&lt;/li&gt;
&lt;li&gt;Do you volunteer the &lt;strong&gt;double-counting trap&lt;/strong&gt; and the &lt;strong&gt;weighting factor&lt;/strong&gt; that fixes it before being asked? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you distinguish a &lt;strong&gt;bridge&lt;/strong&gt; (resolves M:N, no measures) from a &lt;strong&gt;many-to-many fact&lt;/strong&gt; (has its own measures) and a &lt;strong&gt;junk dimension&lt;/strong&gt; (collapses low-cardinality flags)? — vocabulary that separates levels.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — the moment a single foreign key runs out of room
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The cleanest way to feel why bridges exist is to try to model a joint account &lt;em&gt;without&lt;/em&gt; one and watch the model fail. We have three accounts and three customers; account A3 is a joint checking account owned by two people. A single &lt;code&gt;customer_key&lt;/code&gt; on the account dimension can hold only one of them, so the model is forced to either drop an owner or invent repeating columns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Three accounts (A1, A2, A3) and three customers (C1 Ada, C2 Linus, C3 Grace). A1 is owned by C1 alone, A2 by C2 alone, and A3 jointly by C1 and C2. Show why a single &lt;code&gt;owner_customer_key&lt;/code&gt; column on &lt;code&gt;dim_account&lt;/code&gt; cannot represent this, and state the grain of the table that can.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;account&lt;/th&gt;
&lt;th&gt;type&lt;/th&gt;
&lt;th&gt;owners&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A1&lt;/td&gt;
&lt;td&gt;checking&lt;/td&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A2&lt;/td&gt;
&lt;td&gt;savings&lt;/td&gt;
&lt;td&gt;C2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A3&lt;/td&gt;
&lt;td&gt;joint checking&lt;/td&gt;
&lt;td&gt;C1, C2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- ATTEMPT (broken): one owner column on the account dimension&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;dim_account&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;account_key&lt;/span&gt;      &lt;span class="nb"&gt;INT&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;account_type&lt;/span&gt;     &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;owner_customer_key&lt;/span&gt; &lt;span class="nb"&gt;INT&lt;/span&gt;   &lt;span class="c1"&gt;-- can hold ONE value; A3 has two owners&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- The fix: an associative bridge at the grain of the relationship&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;bridge_account_customer&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;account_key&lt;/span&gt;   &lt;span class="nb"&gt;INT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;customer_key&lt;/span&gt;  &lt;span class="nb"&gt;INT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;-- one row per (account, customer) pair&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;bridge_account_customer&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;VALUES&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;           &lt;span class="c1"&gt;-- A1 -&amp;gt; C1&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;           &lt;span class="c1"&gt;-- A2 -&amp;gt; C2&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;   &lt;span class="c1"&gt;-- A3 -&amp;gt; C1 and C2  (the joint account)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The &lt;code&gt;dim_account.owner_customer_key&lt;/code&gt; column is single-valued, so A3 can name only one of its two owners — the model is already lossy before a single query runs. Moving ownership out to &lt;code&gt;bridge_account_customer&lt;/code&gt; changes the grain to "one row per account-customer pair," so A3 simply gets two rows. The composite primary key &lt;code&gt;(account_key, customer_key)&lt;/code&gt; both enforces that each pair appears once and documents the grain. Nothing about the &lt;em&gt;account&lt;/em&gt; or the &lt;em&gt;customer&lt;/em&gt; is stored here — only the fact that they are related.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;bridge row&lt;/th&gt;
&lt;th&gt;account_key&lt;/th&gt;
&lt;th&gt;customer_key&lt;/th&gt;
&lt;th&gt;meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;A1 owned by C1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;A2 owned by C2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;A3 co-owned by C1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;A3 co-owned by C2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If a foreign-key column would ever need to hold more than one value, the relationship is many-to-many and belongs in a bridge whose grain is "one row per related pair."&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Bridge tables between a fact and a dimension
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The account↔customer bridge — how a fact reaches a many-to-many dimension, and the trap on the way
&lt;/h3&gt;

&lt;p&gt;Once the bridge exists, the interesting question is how a &lt;strong&gt;fact&lt;/strong&gt; measured at one grain (a daily account balance) is reported along a dimension it relates to &lt;em&gt;many-to-many&lt;/em&gt; (the customers who own the account). The fact stays at its natural grain — one row per account per day — and the customer is reached by hopping through the bridge. Get the join path right and you can slice balance by customer; get it wrong and the bank's numbers inflate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The join path, one hop at a time.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fact → account.&lt;/strong&gt; &lt;code&gt;fact_balance&lt;/code&gt; carries &lt;code&gt;account_key&lt;/code&gt; at its own grain (account × day). This is a clean many-to-one join, exactly as in any star schema.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Account → bridge.&lt;/strong&gt; &lt;code&gt;bridge_account_customer&lt;/code&gt; explodes each account into one row per owner. This is where a single fact row can fan into several.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bridge → customer.&lt;/strong&gt; The bridge's &lt;code&gt;customer_key&lt;/code&gt; reaches &lt;code&gt;dim_customer&lt;/code&gt; for names, segments, and other attributes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The shape.&lt;/strong&gt; &lt;code&gt;fact → dim_account → bridge → dim_customer&lt;/code&gt; is the canonical four-table path for a fact reported through a many-to-many dimension.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Two places to put the bridge (and why it matters).&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Directly between two dimensions&lt;/strong&gt; — &lt;code&gt;bridge_account_customer&lt;/code&gt; links &lt;code&gt;dim_account&lt;/code&gt; and &lt;code&gt;dim_customer&lt;/code&gt;. Simple, and the pattern shown here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Between a dimension and a "group"&lt;/strong&gt; — Kimball's multivalued-dimension bridge: the account points at a &lt;code&gt;customer_group_key&lt;/code&gt;, and the bridge maps each group to its members. This lets many accounts that share the exact same owner set reuse one group, shrinking the bridge — worth it when groups repeat heavily.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The rule.&lt;/strong&gt; Use a direct bridge when pairs are mostly unique; introduce a group key when the same set of members recurs across many parents.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The double-counting trap, stated precisely.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A fan-out join multiplies the measure.&lt;/strong&gt; Joining &lt;code&gt;fact_balance&lt;/code&gt; through the bridge turns A3's one balance row into two rows (one per owner). &lt;code&gt;SUM(balance)&lt;/code&gt; across the exploded result now counts A3's balance twice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is grain, not a bug in SQL.&lt;/strong&gt; The join is doing exactly what you asked; the result set's grain became "balance per account-owner," so a naive total over it is meaningless.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When it is actually fine.&lt;/strong&gt; A per-customer &lt;em&gt;coverage&lt;/em&gt; or &lt;em&gt;impact&lt;/em&gt; report — "which customers are exposed to this account" — deliberately wants the full value against each owner. The trap only bites when you then sum across customers expecting the bank total.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8ljza9zrpyg8t06brchx.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8ljza9zrpyg8t06brchx.jpeg" alt="Iconographic bridge-table diagram — a balance fact joined to an account dimension, an account-customer bridge in the middle resolving the many-to-many, and a customer dimension on the right, with the joint account A3 fanning out to two owners." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Worked example — slicing balance by customer through the bridge
&lt;/h4&gt;

&lt;p&gt;*&lt;em&gt;Detailed explanation. *&lt;/em&gt; The everyday query is "total balance held by each customer." It reads naturally — join the fact to the account, hop the bridge to the customer, group by customer — and it is correct &lt;em&gt;per customer&lt;/em&gt;. The subtlety is only in what happens when you sum those per-customer numbers back up, which the next section fixes with a weight. Here we establish the honest per-customer view and expose the inflated grand total.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; With balances A1 = 1000, A2 = 4000, A3 = 600, report total balance per customer by joining &lt;code&gt;fact_balance&lt;/code&gt; through the bridge. Then sum those per-customer totals and compare to the true bank total of 5600.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;account_key&lt;/th&gt;
&lt;th&gt;balance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 (A1)&lt;/td&gt;
&lt;td&gt;1000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 (A2)&lt;/td&gt;
&lt;td&gt;4000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 (A3)&lt;/td&gt;
&lt;td&gt;600&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;balance&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;balance_by_customer&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;fact_balance&lt;/span&gt;          &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;bridge_account_customer&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;dim_customer&lt;/span&gt;          &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The join &lt;code&gt;fact_balance → bridge&lt;/code&gt; fans A3's single 600 row into two rows — one for C1, one for C2. C1 therefore accumulates A1 (1000) + A3 (600) = 1600, and C2 accumulates A2 (4000) + A3 (600) = 4600. Both per-customer numbers are legitimate: C1 really does have access to 1600 of balance. But summing the customer column gives 1600 + 4600 = 6200, which overshoots the real 5600 by exactly A3's 600 — the joint balance counted once per owner.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;accounts summed&lt;/th&gt;
&lt;th&gt;balance_by_customer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;C1 Ada&lt;/td&gt;
&lt;td&gt;A1 + A3&lt;/td&gt;
&lt;td&gt;1600&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C2 Linus&lt;/td&gt;
&lt;td&gt;A2 + A3&lt;/td&gt;
&lt;td&gt;4600&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;grand total (naive)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;6200 ✗ (true = 5600)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; A fact summed &lt;em&gt;through&lt;/em&gt; a bridge is correct along the grouped dimension but not across it — the per-customer numbers are trustworthy, the grand total is inflated by every shared parent.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on a fact-to-dimension bridge
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You are handed &lt;code&gt;fact_balance(account_key, balance)&lt;/code&gt;, &lt;code&gt;bridge_account_customer(account_key, customer_key)&lt;/code&gt;, and &lt;code&gt;dim_customer(customer_key, name)&lt;/code&gt;. An analyst wrote &lt;code&gt;SELECT SUM(balance) FROM fact_balance JOIN bridge_account_customer USING(account_key)&lt;/code&gt; to get total bank balance and got 6200 instead of 5600. Explain the defect and write a query that returns the correct bank total while still supporting a per-customer breakdown.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using a pre-aggregated fact to protect the grain
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Correct bank total: aggregate the fact at its OWN grain, never through the bridge&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;balance&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;bank_total&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_balance&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;                     &lt;span class="c1"&gt;-- 1000 + 4000 + 600 = 5600&lt;/span&gt;

&lt;span class="c1"&gt;-- Per-customer breakdown: fan through the bridge on purpose, but label it "exposure"&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;balance&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;customer_exposure&lt;/span&gt;   &lt;span class="c1"&gt;-- intentionally counts shared balance per owner&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;fact_balance&lt;/span&gt;           &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;bridge_account_customer&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;dim_customer&lt;/span&gt;           &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;rows&lt;/th&gt;
&lt;th&gt;balance seen&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1. fact at native grain&lt;/td&gt;
&lt;td&gt;A1, A2, A3&lt;/td&gt;
&lt;td&gt;1000, 4000, 600&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2. SUM without joining bridge&lt;/td&gt;
&lt;td&gt;3 rows&lt;/td&gt;
&lt;td&gt;5600 (correct total)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3. join bridge (exposure query)&lt;/td&gt;
&lt;td&gt;A1→C1, A2→C2, A3→C1, A3→C2&lt;/td&gt;
&lt;td&gt;4 rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4. group by customer&lt;/td&gt;
&lt;td&gt;C1=1600, C2=4600&lt;/td&gt;
&lt;td&gt;per-customer exposure&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The bank total must be computed &lt;strong&gt;before&lt;/strong&gt; the bridge fans the fact out; &lt;code&gt;SUM(balance)&lt;/code&gt; over &lt;code&gt;fact_balance&lt;/code&gt; alone touches each account once, so A3's 600 is counted once → 5600.&lt;/li&gt;
&lt;li&gt;The moment the bridge joins in, A3 becomes two rows, so any total over that result double-counts by design.&lt;/li&gt;
&lt;li&gt;The per-customer query keeps the fan-out, but you &lt;em&gt;name&lt;/em&gt; the measure &lt;code&gt;customer_exposure&lt;/code&gt; so no one mistakes it for an additive, cross-customer total.&lt;/li&gt;
&lt;li&gt;Two questions, two grains, two queries — the fix is recognizing they are different questions, not forcing one query to answer both.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;th&gt;value&lt;/th&gt;
&lt;th&gt;grain&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;bank_total&lt;/td&gt;
&lt;td&gt;5600&lt;/td&gt;
&lt;td&gt;fact grain (correct)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;exposure C1&lt;/td&gt;
&lt;td&gt;1600&lt;/td&gt;
&lt;td&gt;per account-owner&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;exposure C2&lt;/td&gt;
&lt;td&gt;4600&lt;/td&gt;
&lt;td&gt;per account-owner&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Grain discipline&lt;/strong&gt;&lt;/strong&gt; — the fact's additive grain is "one row per account"; any &lt;code&gt;SUM&lt;/code&gt; that must equal the true total has to be taken at that grain, before a bridge multiplies rows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Fan-out is intentional&lt;/strong&gt;&lt;/strong&gt; — joining through the bridge is the right move for a per-customer exposure view; it is only wrong when you then aggregate across the fanned dimension.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Name the measure by its grain&lt;/strong&gt;&lt;/strong&gt; — calling the fanned number &lt;code&gt;customer_exposure&lt;/code&gt; rather than &lt;code&gt;balance&lt;/code&gt; stops a downstream consumer from summing it into a meaningless total.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Two questions, two queries&lt;/strong&gt;&lt;/strong&gt; — "what is the bank total" and "what is each customer exposed to" have different grains and deserve separate SQL rather than one query patched to fake both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the total is O(fact rows); the exposure query is O(fact rows × avg owners per account), bounded by the bridge fan-out factor.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;SQL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — bridge-table&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Bridge-table join and grain problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/bridge-table" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Modeling&lt;/span&gt;
&lt;span&gt;Topic — many-to-many&lt;/span&gt;
&lt;strong&gt;Many-to-many relationship modeling problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/many-to-many" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. Weighting factors — allocation without double-counting
&lt;/h2&gt;
&lt;h3&gt;
  
  
  An allocation factor makes a shared measure sum correctly — one column that turns a fan trap into a clean total
&lt;/h3&gt;

&lt;p&gt;The fix for the double-counting you just saw is a single extra column on the bridge: a &lt;strong&gt;weighting factor&lt;/strong&gt; (also called an allocation factor) that says what fraction of a parent's measure belongs to each related member. Store weights that &lt;strong&gt;sum to 1.0 per parent&lt;/strong&gt;, multiply the measure by the weight in your aggregate, and the shared value is split instead of duplicated — so a cross-member total lands exactly on the true whole.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The mechanism.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A weight per bridge row.&lt;/strong&gt; &lt;code&gt;bridge_account_customer&lt;/code&gt; gains a &lt;code&gt;weighting_factor&lt;/code&gt; column. For a sole-owner account it is 1.0; for a 50/50 joint account each owner's row carries 0.5.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The invariant.&lt;/strong&gt; For any given parent (account), the weights across all its bridge rows &lt;strong&gt;sum to exactly 1.0&lt;/strong&gt;. This is the property that guarantees the allocated total equals the original.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The corrected expression.&lt;/strong&gt; Replace &lt;code&gt;SUM(f.balance)&lt;/code&gt; with &lt;code&gt;SUM(f.balance * b.weighting_factor)&lt;/code&gt;. Now A3's 600 contributes 300 to C1 and 300 to C2 — 600 in total, not 1200.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where weights come from.&lt;/strong&gt; Ownership percentages, headcount splits, revenue-share agreements, or a default equal split (&lt;code&gt;1.0 / number_of_owners&lt;/code&gt;) when no business rule exists.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Two legitimate reports, two different measures.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Impact / coverage report (unweighted).&lt;/strong&gt; Uses the raw measure against every member — "how much balance does each customer have access to." Correct per member; must not be summed across members.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Allocated report (weighted).&lt;/strong&gt; Uses &lt;code&gt;measure × weight&lt;/code&gt; — "how much of the balance is attributable to each customer." Sums correctly to the whole and is the one finance signs off on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ship both, label both.&lt;/strong&gt; Interviewers love that you keep them distinct: an impact number and an allocated number answer different business questions and both are valid.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Traps around weighting.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Weights that don't sum to 1.0.&lt;/strong&gt; If a parent's weights sum to 0.9 or 1.1, your allocated total drifts below or above the truth; enforce the invariant with a validation query or a CHECK on a materialized rollup.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weighting a non-additive measure.&lt;/strong&gt; Allocation is meaningful for additive measures (balance, revenue). A weighted average of a ratio (e.g. interest rate) needs a weighted-average formula, not a straight &lt;code&gt;measure × weight&lt;/code&gt; sum.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Missing bridge rows.&lt;/strong&gt; An account with &lt;em&gt;no&lt;/em&gt; bridge rows drops out of every through-the-bridge query entirely; guard with an outer join or a data-quality check for orphan accounts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdlu4ph69ypo639hcynnc.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdlu4ph69ypo639hcynnc.jpeg" alt="Iconographic allocation-factor diagram — a joint account balance split 50/50 across two owners by a weighting factor, contrasting an unweighted impact report that double-counts with a weighted report that sums correctly." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — the same query, corrected by a weight
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; We take the exact per-customer query from section 2 and add &lt;code&gt;* b.weighting_factor&lt;/code&gt;. A3's 600 now splits 300/300 across its two owners, so the per-customer numbers change and — crucially — their sum collapses onto the real 5600. This is the whole trick: one column, one multiplication.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Add a &lt;code&gt;weighting_factor&lt;/code&gt; to the bridge (sole accounts 1.0; A3 split 0.5/0.5) and report allocated balance per customer. Confirm the allocated total equals the true bank total of 5600.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;account_key&lt;/th&gt;
&lt;th&gt;customer_key&lt;/th&gt;
&lt;th&gt;weighting_factor&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 (A1)&lt;/td&gt;
&lt;td&gt;1 (C1)&lt;/td&gt;
&lt;td&gt;1.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 (A2)&lt;/td&gt;
&lt;td&gt;2 (C2)&lt;/td&gt;
&lt;td&gt;1.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 (A3)&lt;/td&gt;
&lt;td&gt;1 (C1)&lt;/td&gt;
&lt;td&gt;0.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 (A3)&lt;/td&gt;
&lt;td&gt;2 (C2)&lt;/td&gt;
&lt;td&gt;0.5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;balance&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;weighting_factor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;allocated_balance&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;fact_balance&lt;/span&gt;            &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;bridge_account_customer&lt;/span&gt;  &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;dim_customer&lt;/span&gt;             &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; For C1: A1 contributes 1000 × 1.0 = 1000 and A3 contributes 600 × 0.5 = 300, so C1 = 1300. For C2: A2 contributes 4000 × 1.0 = 4000 and A3 contributes 600 × 0.5 = 300, so C2 = 4300. The joint account's 600 is now split rather than duplicated, so 1300 + 4300 = 5600 — exactly the bank total. The only change from the section-2 query is the &lt;code&gt;* b.weighting_factor&lt;/code&gt; factor inside the &lt;code&gt;SUM&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer&lt;/th&gt;
&lt;th&gt;allocated_balance&lt;/th&gt;
&lt;th&gt;contribution from A3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;C1 Ada&lt;/td&gt;
&lt;td&gt;1300&lt;/td&gt;
&lt;td&gt;300&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C2 Linus&lt;/td&gt;
&lt;td&gt;4300&lt;/td&gt;
&lt;td&gt;300&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;allocated total&lt;/td&gt;
&lt;td&gt;5600 ✓&lt;/td&gt;
&lt;td&gt;600 (split, not doubled)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; When weights on a bridge sum to 1.0 per parent, &lt;code&gt;SUM(measure × weight)&lt;/code&gt; is safe to total across the many-to-many dimension; without the weight, it never is.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on allocation
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Marketing wants "attributed revenue per customer" from &lt;code&gt;fact_revenue(account_key, revenue)&lt;/code&gt; through the account↔customer bridge, and the CFO insists the per-customer numbers must add up to total revenue. Some accounts are joint. Write the query, and describe how you would validate that the weighting factors are trustworthy.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using weighted allocation with a weight-integrity check
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Attributed revenue that sums to the whole&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;revenue&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;weighting_factor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;attributed_revenue&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;fact_revenue&lt;/span&gt;            &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;bridge_account_customer&lt;/span&gt;  &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;dim_customer&lt;/span&gt;             &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- Weight-integrity check: every account's weights must sum to 1.0&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="n"&gt;account_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;weighting_factor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;weight_sum&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;bridge_account_customer&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="n"&gt;account_key&lt;/span&gt;
&lt;span class="k"&gt;HAVING&lt;/span&gt;      &lt;span class="k"&gt;ABS&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;weighting_factor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0001&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;-- rows here are BAD&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;account&lt;/th&gt;
&lt;th&gt;owners&lt;/th&gt;
&lt;th&gt;weights&lt;/th&gt;
&lt;th&gt;revenue&lt;/th&gt;
&lt;th&gt;contribution split&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A1&lt;/td&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;1.0&lt;/td&gt;
&lt;td&gt;1000&lt;/td&gt;
&lt;td&gt;C1 += 1000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A2&lt;/td&gt;
&lt;td&gt;C2&lt;/td&gt;
&lt;td&gt;1.0&lt;/td&gt;
&lt;td&gt;4000&lt;/td&gt;
&lt;td&gt;C2 += 4000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A3&lt;/td&gt;
&lt;td&gt;C1, C2&lt;/td&gt;
&lt;td&gt;0.5, 0.5&lt;/td&gt;
&lt;td&gt;600&lt;/td&gt;
&lt;td&gt;C1 += 300, C2 += 300&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The allocation query multiplies each account's revenue by the owner's &lt;code&gt;weighting_factor&lt;/code&gt;, so a joint account's revenue is divided among owners instead of replicated.&lt;/li&gt;
&lt;li&gt;Because weights per account sum to 1.0, the sum of all &lt;code&gt;attributed_revenue&lt;/code&gt; equals &lt;code&gt;SUM(revenue)&lt;/code&gt; over the fact — the CFO's invariant holds by construction.&lt;/li&gt;
&lt;li&gt;The integrity check groups the bridge by &lt;code&gt;account_key&lt;/code&gt; and flags any account whose weights do &lt;strong&gt;not&lt;/strong&gt; sum to 1.0 (within a tiny epsilon for floating point); an empty result means the weights are trustworthy.&lt;/li&gt;
&lt;li&gt;Run the check as a data-quality test in CI or the pipeline so a bad weight fails the build rather than silently skewing attributed revenue.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;th&gt;value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;attributed C1&lt;/td&gt;
&lt;td&gt;1300&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;attributed C2&lt;/td&gt;
&lt;td&gt;4300&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sum attributed&lt;/td&gt;
&lt;td&gt;5600 (= total revenue)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;integrity check rows&lt;/td&gt;
&lt;td&gt;0 (all accounts sum to 1.0)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Weighting factor&lt;/strong&gt;&lt;/strong&gt; — a per-relationship fraction stored on the bridge; multiplying the measure by it converts duplication into division.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Sum-to-one invariant&lt;/strong&gt;&lt;/strong&gt; — enforcing that each parent's weights total 1.0 is precisely what makes the allocated grand total equal the un-allocated one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Allocation vs impact&lt;/strong&gt;&lt;/strong&gt; — weighted gives attributable revenue that adds up; unweighted gives exposure per member; naming them apart prevents a wrong total.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Weight integrity as a test&lt;/strong&gt;&lt;/strong&gt; — a &lt;code&gt;HAVING SUM(weight) &amp;lt;&amp;gt; 1.0&lt;/code&gt; query turns a subtle correctness property into an automatable data-quality gate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — one extra multiply per bridged row, O(fact rows × fan-out); the integrity check is O(bridge rows) and runs once per load, not per query.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Allocation&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — allocation-factor&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Allocation-factor and weighted-rollup problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/allocation-factor" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;SQL&lt;/span&gt;
&lt;span&gt;Topic — bridge-table&lt;/span&gt;
&lt;strong&gt;Bridge-table double-counting problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/bridge-table" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Multi-valued attributes &amp;amp; multi-valued dimensions
&lt;/h2&gt;
&lt;h3&gt;
  
  
  One member, many values — a multi-valued dimension attaches several attribute values through a bridge
&lt;/h3&gt;

&lt;p&gt;The other half of the bridge story is the &lt;strong&gt;multi-valued dimension&lt;/strong&gt;: a single dimension member that legitimately carries several values of one attribute. A customer belongs to several marketing segments; a patient visit carries several diagnoses; a product wears several category tags; a job posting lists several required skills. The wrong instinct is a comma-separated string or a fixed set of &lt;code&gt;segment_1..segment_n&lt;/code&gt; columns; the right pattern is a bridge from the member to a value dimension.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The shape of a multi-valued dimension.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Member → bridge → value dimension.&lt;/strong&gt; &lt;code&gt;dim_customer → bridge_customer_segment → dim_segment&lt;/code&gt;. The bridge holds one row per &lt;code&gt;(customer_key, segment_key)&lt;/code&gt; pair — the same associative grain as before, now between a dimension and its multi-valued attribute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optional group dimension.&lt;/strong&gt; When many members share the exact same set of values (thousands of customers all tagged &lt;code&gt;{premium, digital}&lt;/code&gt;), insert a &lt;code&gt;segment_group&lt;/code&gt; between the member and the bridge so identical sets are stored once. The customer points at a &lt;code&gt;segment_group_key&lt;/code&gt;; the bridge maps group → segments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optional weight.&lt;/strong&gt; Just like ownership, a multi-valued attribute can carry a weight (primary vs secondary diagnosis, a confidence score), reusing the allocation machinery from section 3.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Querying a multi-valued dimension correctly.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Filtering is easy and safe.&lt;/strong&gt; "Customers in the &lt;code&gt;premium&lt;/code&gt; segment" is a straightforward join &lt;code&gt;dim_customer → bridge → dim_segment WHERE segment = 'premium'&lt;/code&gt;; filtering does not double-count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Counting members needs &lt;code&gt;COUNT(DISTINCT)&lt;/code&gt;.&lt;/strong&gt; "How many customers per segment" fans a customer into one row per segment; counting &lt;em&gt;customers&lt;/em&gt; across all segments must use &lt;code&gt;COUNT(DISTINCT customer_key)&lt;/code&gt; or you count a two-segment customer twice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Summing a member's measure needs a weight.&lt;/strong&gt; If you attach a customer-level measure (lifetime value) and report it by segment, you are back in double-counting territory — allocate with a weight exactly as in section 3.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Multi-valued attribute vs the alternatives.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;vs comma-list.&lt;/strong&gt; A &lt;code&gt;'premium,digital'&lt;/code&gt; string can't be joined, indexed, or filtered without fragile string matching; a bridge makes every value a first-class, queryable row.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vs pivoted columns.&lt;/strong&gt; &lt;code&gt;is_premium&lt;/code&gt;, &lt;code&gt;is_digital&lt;/code&gt; boolean columns work only for a small, fixed, rarely-changing value set; a new segment means a schema change. A bridge scales to any number of values with no DDL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vs a junk dimension.&lt;/strong&gt; A junk dimension collapses several &lt;em&gt;low-cardinality, unrelated&lt;/em&gt; flags into one row; it is not for a single attribute with many values that relate a member to a growing list. Different tool, different problem.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnhyldc1nmkgivewvek32.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnhyldc1nmkgivewvek32.jpeg" alt="Iconographic multi-valued dimension diagram — one customer carrying several segment values through a bridge to a segment dimension, with a group dimension deduplicating common value sets." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — customers with several segments
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; We tag customers with marketing segments through a bridge. C1 Ada is &lt;code&gt;{premium, digital}&lt;/code&gt;, C2 Linus is &lt;code&gt;{retail}&lt;/code&gt;, C3 Grace is &lt;code&gt;{premium, retail}&lt;/code&gt;. The example shows that filtering by segment is safe, but counting customers across segments needs &lt;code&gt;COUNT(DISTINCT)&lt;/code&gt; because a multi-segment customer appears in several bridge rows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Given the segment tags above, (a) list customers in the &lt;code&gt;premium&lt;/code&gt; segment, and (b) count distinct customers per segment. Show why a plain &lt;code&gt;COUNT(*)&lt;/code&gt; on a cross-segment total would over-count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;customer_key&lt;/th&gt;
&lt;th&gt;segment_key&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 (C1)&lt;/td&gt;
&lt;td&gt;premium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1 (C1)&lt;/td&gt;
&lt;td&gt;digital&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 (C2)&lt;/td&gt;
&lt;td&gt;retail&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 (C3)&lt;/td&gt;
&lt;td&gt;premium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 (C3)&lt;/td&gt;
&lt;td&gt;retail&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- (a) customers in the 'premium' segment  (filtering is safe)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;dim_customer&lt;/span&gt;          &lt;span class="k"&gt;c&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;bridge_customer_segment&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;dim_segment&lt;/span&gt;           &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;       &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'premium'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- (b) distinct customers per segment  (COUNT DISTINCT guards the fan-out)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;bridge_customer_segment&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;dim_segment&lt;/span&gt;           &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; Query (a) filters the bridge to &lt;code&gt;premium&lt;/code&gt; rows and returns C1 and C3 — filtering never duplicates because we are selecting members, not aggregating a measure across the fan-out. Query (b) groups by segment; &lt;code&gt;premium&lt;/code&gt; has two rows (C1, C3), &lt;code&gt;retail&lt;/code&gt; has two (C2, C3), &lt;code&gt;digital&lt;/code&gt; has one (C1). &lt;code&gt;COUNT(DISTINCT customer_key)&lt;/code&gt; counts members correctly within each segment. If you instead summed those segment counts to get "total customers" you would get 2 + 2 + 1 = 5 for only 3 real customers — because C1 and C3 each sit in two segments — which is why a cross-segment headcount must &lt;code&gt;COUNT(DISTINCT customer_key)&lt;/code&gt; over the whole set, not add the per-segment counts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;segment&lt;/th&gt;
&lt;th&gt;customers&lt;/th&gt;
&lt;th&gt;members&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;premium&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;C1, C3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;digital&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;retail&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;C2, C3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Filtering a multi-valued dimension is free; counting members across it always needs &lt;code&gt;COUNT(DISTINCT)&lt;/code&gt;, and summing a member measure across it always needs a weight.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on a multi-valued dimension
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Each product has one or more category tags via &lt;code&gt;bridge_product_category&lt;/code&gt;, and &lt;code&gt;fact_sales(product_key, revenue)&lt;/code&gt; records sales at the product grain. A PM asks for "revenue by category." Some products have several categories. Give the report that (a) shows revenue exposure per category and (b) also produces category revenue that sums to total sales, and explain the difference.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using COUNT(DISTINCT) plus weighted category allocation
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- (a) exposure: full product revenue against every category it belongs to&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="n"&gt;cat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;revenue&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;revenue_exposure&lt;/span&gt;     &lt;span class="c1"&gt;-- sums &amp;gt; total sales if products are multi-category&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;fact_sales&lt;/span&gt;            &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;bridge_product_category&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;dim_category&lt;/span&gt;          &lt;span class="n"&gt;cat&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;cat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="n"&gt;cat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category_name&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- (b) allocated: split each product's revenue evenly across its categories&lt;/span&gt;
&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="n"&gt;cat_counts&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;product_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;n_categories&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;bridge_product_category&lt;/span&gt;
    &lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;product_key&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="n"&gt;cat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;revenue&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;cc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;n_categories&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;allocated_revenue&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;fact_sales&lt;/span&gt;            &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;bridge_product_category&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;cat_counts&lt;/span&gt;            &lt;span class="n"&gt;cc&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;cc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;dim_category&lt;/span&gt;          &lt;span class="n"&gt;cat&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;cat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="n"&gt;cat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category_name&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;product&lt;/th&gt;
&lt;th&gt;categories&lt;/th&gt;
&lt;th&gt;revenue&lt;/th&gt;
&lt;th&gt;exposure adds&lt;/th&gt;
&lt;th&gt;allocated adds&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;P1&lt;/td&gt;
&lt;td&gt;electronics&lt;/td&gt;
&lt;td&gt;1000&lt;/td&gt;
&lt;td&gt;electronics += 1000&lt;/td&gt;
&lt;td&gt;electronics += 1000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;P2&lt;/td&gt;
&lt;td&gt;electronics, home&lt;/td&gt;
&lt;td&gt;600&lt;/td&gt;
&lt;td&gt;each += 600&lt;/td&gt;
&lt;td&gt;each += 300&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The exposure query (a) sums full product revenue against every category, so P2's 600 lands in both &lt;code&gt;electronics&lt;/code&gt; and &lt;code&gt;home&lt;/code&gt; — great for "which categories touch this revenue," but the category column now totals more than sales.&lt;/li&gt;
&lt;li&gt;The allocated query (b) derives &lt;code&gt;n_categories&lt;/code&gt; per product in a CTE, then divides revenue by that count — an equal-split weighting factor computed on the fly (&lt;code&gt;1.0 / n_categories&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Because each product's split weights sum to 1.0, the allocated category revenue sums back to total sales exactly, satisfying "must add up."&lt;/li&gt;
&lt;li&gt;The two answers are both correct for different questions: exposure measures reach, allocation measures attributable revenue.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;category&lt;/th&gt;
&lt;th&gt;revenue_exposure (a)&lt;/th&gt;
&lt;th&gt;allocated_revenue (b)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;electronics&lt;/td&gt;
&lt;td&gt;1600&lt;/td&gt;
&lt;td&gt;1300&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;home&lt;/td&gt;
&lt;td&gt;600&lt;/td&gt;
&lt;td&gt;300&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;column total&lt;/td&gt;
&lt;td&gt;2200 (&amp;gt; 1600 sales)&lt;/td&gt;
&lt;td&gt;1600 (= total sales)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Multi-valued dimension&lt;/strong&gt;&lt;/strong&gt; — modeling category as a bridge, not a column, lets a product carry any number of categories and keeps every one queryable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Exposure vs allocation&lt;/strong&gt;&lt;/strong&gt; — exposure keeps the full measure per value (reach); allocation splits it (attribution); each answers a real, distinct question.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;On-the-fly weight&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;1.0 / n_categories&lt;/code&gt; is a default equal-split weighting factor derived when the bridge has no stored weight, reusing the section-3 machinery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;DISTINCT for headcounts&lt;/strong&gt;&lt;/strong&gt; — whenever the deliverable is a count of members rather than a measure, &lt;code&gt;COUNT(DISTINCT member_key)&lt;/code&gt; is the fan-out guard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the CTE is O(bridge rows) to compute counts; the allocation join is O(fact rows × avg categories per product).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Modeling&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — many-to-many&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Multi-valued dimension modeling problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/many-to-many" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;SQL&lt;/span&gt;
&lt;span&gt;Topic — bridge-table&lt;/span&gt;
&lt;strong&gt;Bridge-table aggregation problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/bridge-table" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Hierarchy bridges &amp;amp; closure tables for ragged trees
&lt;/h2&gt;
&lt;h3&gt;
  
  
  A closure table stores every ancestor-descendant pair — one join rolls up a fact at any level of a ragged tree
&lt;/h3&gt;

&lt;p&gt;The hardest many-to-many is a member relating to itself at unknown depth: an org chart, a bill of materials, a category tree. An &lt;strong&gt;adjacency list&lt;/strong&gt; (&lt;code&gt;employee.manager_key&lt;/code&gt;) captures each parent-child edge but can't be flattened by a fixed number of joins because the depth is variable — a &lt;strong&gt;ragged hierarchy&lt;/strong&gt;. The bridge that solves this is a &lt;strong&gt;closure table&lt;/strong&gt; (Kimball calls it a hierarchy bridge): one row for every ancestor-descendant pair, including each node to itself, tagged with the depth between them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What a closure table holds.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One row per ancestor→descendant pair.&lt;/strong&gt; For every node, a row to each of its descendants at any depth, plus a depth-0 self row. E1's rows reach E1 (0), E2 (1), E3 (1), E4 (2), E5 (2).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A depth column.&lt;/strong&gt; &lt;code&gt;depth&lt;/code&gt; (a.k.a. distance) is the number of edges between ancestor and descendant; the self row has depth 0. It lets you query "direct reports only" (&lt;code&gt;depth = 1&lt;/code&gt;) or "everyone below" (&lt;code&gt;depth &amp;gt;= 1&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optional flags.&lt;/strong&gt; &lt;code&gt;is_leaf&lt;/code&gt;, &lt;code&gt;is_root&lt;/code&gt;, or Kimball's &lt;code&gt;lowest_flag&lt;/code&gt; mark where a node sits, which is handy for stopping double-counts when a fact can attach at multiple levels.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No measures.&lt;/strong&gt; Like every bridge, it stores structure, not facts; you join facts &lt;em&gt;through&lt;/em&gt; it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Building and using it.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Build with a recursive CTE.&lt;/strong&gt; Walk the adjacency list from every node downward, accumulating depth; the anchor emits self rows at depth 0 and the recursive member extends each path by one edge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Roll up a subtree with one join.&lt;/strong&gt; To total a measure for "everyone under E2," join the fact on &lt;code&gt;descendant&lt;/code&gt; and filter &lt;code&gt;ancestor = E2&lt;/code&gt;. One join replaces an unbounded chain of self-joins.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Direct vs full rollup by depth.&lt;/strong&gt; &lt;code&gt;WHERE ancestor = X AND depth = 1&lt;/code&gt; gives direct children; dropping the depth filter (or &lt;code&gt;depth &amp;gt;= 0&lt;/code&gt;) gives the whole subtree including X itself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Traps in hierarchy rollups.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Double-counting across ancestors.&lt;/strong&gt; Every node has multiple ancestors (itself, its manager, its manager's manager). If you join a fact to the closure table and forget to fix a single &lt;code&gt;ancestor&lt;/code&gt;, each fact row is counted once per ancestor — a massive over-count. Always filter or group by one ancestor level.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Facts attached at inner nodes.&lt;/strong&gt; If measures live only at leaves, a plain descendant rollup is clean; if inner nodes also carry measures (a manager has personal sales), the depth-0 self rows correctly include them — which is exactly why the self row exists.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cycles.&lt;/strong&gt; A malformed hierarchy with a cycle makes the recursive CTE loop; guard with a depth cap or a visited-path check, and validate that the source is a true tree/DAG.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgoilzesu45440hhmcwja.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgoilzesu45440hhmcwja.jpeg" alt="Iconographic closure-table diagram — a ragged org-chart tree on the left and its ancestor-descendant closure table on the right with depth values, used to roll up a sales fact for a whole subtree with one join." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — building a closure table from an org chart
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; We have a five-person org: E1 (CEO) manages E2 and E3; E2 manages E4 and E5. The adjacency list stores only the direct &lt;code&gt;manager_key&lt;/code&gt;. A recursive CTE expands it into a closure table of every ancestor-descendant pair with depth, which is the structure we then roll facts up through.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; From &lt;code&gt;dim_employee(employee_key, manager_key)&lt;/code&gt; with edges E1→E2, E1→E3, E2→E4, E2→E5, build the closure table of &lt;code&gt;(ancestor, descendant, depth)&lt;/code&gt; including depth-0 self rows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;employee_key&lt;/th&gt;
&lt;th&gt;manager_key&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;E1&lt;/td&gt;
&lt;td&gt;(none)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E2&lt;/td&gt;
&lt;td&gt;E1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E3&lt;/td&gt;
&lt;td&gt;E1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E4&lt;/td&gt;
&lt;td&gt;E2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E5&lt;/td&gt;
&lt;td&gt;E2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="k"&gt;RECURSIVE&lt;/span&gt; &lt;span class="n"&gt;closure&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="c1"&gt;-- anchor: every node is its own ancestor at depth 0&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="n"&gt;employee_key&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;ancestor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;employee_key&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;descendant&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="mi"&gt;0&lt;/span&gt;            &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;depth&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt;    &lt;span class="n"&gt;dim_employee&lt;/span&gt;

    &lt;span class="k"&gt;UNION&lt;/span&gt; &lt;span class="k"&gt;ALL&lt;/span&gt;

    &lt;span class="c1"&gt;-- recursive: extend each known pair down one edge&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ancestor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;employee_key&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;descendant&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;depth&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;    &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;depth&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;closure&lt;/span&gt;       &lt;span class="k"&gt;c&lt;/span&gt;
    &lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;dim_employee&lt;/span&gt;  &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;manager_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;descendant&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;ancestor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;descendant&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;depth&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;closure&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;ancestor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;depth&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;descendant&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The anchor emits five self rows at depth 0 (E1→E1, …, E5→E5). The recursive member joins each existing pair's &lt;code&gt;descendant&lt;/code&gt; to employees who report to it, emitting a new pair with &lt;code&gt;depth + 1&lt;/code&gt;: from E1→E1 it finds E2 and E3 (depth 1), then from E1→E2 it finds E4 and E5 (depth 2), and independently E2→E2 yields E2→E4 and E2→E5 (depth 1). Recursion stops when no node reports to the current descendants (leaves E3, E4, E5). The result is every ancestor-descendant pair with its depth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;ancestor&lt;/th&gt;
&lt;th&gt;descendant&lt;/th&gt;
&lt;th&gt;depth&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;E1&lt;/td&gt;
&lt;td&gt;E1&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E1&lt;/td&gt;
&lt;td&gt;E2&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E1&lt;/td&gt;
&lt;td&gt;E3&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E1&lt;/td&gt;
&lt;td&gt;E4&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E1&lt;/td&gt;
&lt;td&gt;E5&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E2&lt;/td&gt;
&lt;td&gt;E2&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E2&lt;/td&gt;
&lt;td&gt;E4&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E2&lt;/td&gt;
&lt;td&gt;E5&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E3&lt;/td&gt;
&lt;td&gt;E3&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E4&lt;/td&gt;
&lt;td&gt;E4&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E5&lt;/td&gt;
&lt;td&gt;E5&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; A closure table trades storage (roughly nodes × average depth rows) for O(1)-join rollups at any level — build it once per hierarchy load, query it cheaply forever.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on hierarchy rollup
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Using the closure table above and &lt;code&gt;fact_sales(employee_key, sales)&lt;/code&gt; with E1=0, E2=100, E3=200, E4=50, E5=70, return total sales for the entire subtree under a given manager (including the manager's own sales). Show the result for E2, and explain how you avoid counting a person's sales once per ancestor.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using a descendant join filtered to one ancestor
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Subtree rollup: everyone at or below the chosen manager&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;  &lt;span class="n"&gt;cl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ancestor&lt;/span&gt;      &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;manager_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sales&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;subtree_sales&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;        &lt;span class="n"&gt;closure&lt;/span&gt;     &lt;span class="n"&gt;cl&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;        &lt;span class="n"&gt;fact_sales&lt;/span&gt;   &lt;span class="n"&gt;f&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;employee_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;descendant&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;       &lt;span class="n"&gt;cl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ancestor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'E2'&lt;/span&gt;          &lt;span class="c1"&gt;-- one ancestor =&amp;gt; no cross-ancestor double count&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt;    &lt;span class="n"&gt;cl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ancestor&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- Direct reports only would add:  AND cl.depth = 1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;closure row (ancestor=E2)&lt;/th&gt;
&lt;th&gt;descendant&lt;/th&gt;
&lt;th&gt;sales joined&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;E2 → E2 (depth 0)&lt;/td&gt;
&lt;td&gt;E2&lt;/td&gt;
&lt;td&gt;100 (manager's own)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E2 → E4 (depth 1)&lt;/td&gt;
&lt;td&gt;E4&lt;/td&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E2 → E5 (depth 1)&lt;/td&gt;
&lt;td&gt;E5&lt;/td&gt;
&lt;td&gt;70&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Filtering &lt;code&gt;ancestor = 'E2'&lt;/code&gt; selects exactly the rows whose descendants form E2's subtree: E2 (self, depth 0), E4, and E5.&lt;/li&gt;
&lt;li&gt;Joining &lt;code&gt;fact_sales&lt;/code&gt; on &lt;code&gt;descendant&lt;/code&gt; attaches each person's sales to precisely one row, because within a single ancestor each descendant appears once.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;SUM(f.sales)&lt;/code&gt; totals 100 + 50 + 70 = 220 — the manager's own sales plus every report at any depth.&lt;/li&gt;
&lt;li&gt;The double-count trap is avoided by fixing a single &lt;code&gt;ancestor&lt;/code&gt;; had we joined the fact to the &lt;em&gt;whole&lt;/em&gt; closure table and grouped by employee, each person's sales would be added once per ancestor above them (E4's 50 would count under E4, E2, and E1). Pinning one ancestor per query is the discipline.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;manager&lt;/th&gt;
&lt;th&gt;subtree_sales&lt;/th&gt;
&lt;th&gt;includes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;E2&lt;/td&gt;
&lt;td&gt;220&lt;/td&gt;
&lt;td&gt;E2 (100) + E4 (50) + E5 (70)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Closure table&lt;/strong&gt;&lt;/strong&gt; — precomputing every ancestor-descendant pair turns a variable-depth walk into a flat join, so any subtree is reachable in one hop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Depth column&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;depth = 1&lt;/code&gt; isolates direct reports, &lt;code&gt;depth &amp;gt;= 0&lt;/code&gt; takes the full subtree; one table answers both without restructuring.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Self rows&lt;/strong&gt;&lt;/strong&gt; — the depth-0 row is what folds a manager's own measure into the subtree total, so inner-node facts are not lost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;One ancestor per query&lt;/strong&gt;&lt;/strong&gt; — fixing a single &lt;code&gt;ancestor&lt;/code&gt; is the guard that stops a fact from being counted once per level of the tree above it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — rollup is O(descendants of the chosen node), a single indexed join; the build is O(nodes × average depth) done once per load, not per query.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;Hierarchy&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — recursive-hierarchy&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Recursive-hierarchy and org-chart rollup problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/recursive-hierarchy" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Modeling&lt;/span&gt;
&lt;span&gt;Topic — closure-table&lt;/span&gt;
&lt;strong&gt;Closure-table and ragged-hierarchy problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/closure-table" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — bridge &amp;amp; hierarchy recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Bridge table DDL (grain = one row per pair).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;bridge_account_customer&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;account_key&lt;/span&gt;      &lt;span class="nb"&gt;INT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;customer_key&lt;/span&gt;     &lt;span class="nb"&gt;INT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;weighting_factor&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Weighted rollup (sums correctly across the M:N dimension).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;balance&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;weighting_factor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;allocated&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_balance&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;bridge_account_customer&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_key&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_customer&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;            &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Weight-integrity check (each parent's weights sum to 1.0).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;account_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;weighting_factor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;bridge_account_customer&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;account_key&lt;/span&gt;
&lt;span class="k"&gt;HAVING&lt;/span&gt; &lt;span class="k"&gt;ABS&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;weighting_factor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0001&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;-- rows = bad&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Multi-valued dimension count (guard the fan-out).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;bridge_customer_segment&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_segment&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Closure table build (recursive CTE over an adjacency list).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="k"&gt;RECURSIVE&lt;/span&gt; &lt;span class="n"&gt;closure&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;employee_key&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;ancestor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;employee_key&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;descendant&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;depth&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;dim_employee&lt;/span&gt;
    &lt;span class="k"&gt;UNION&lt;/span&gt; &lt;span class="k"&gt;ALL&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ancestor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;employee_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;depth&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;closure&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
    &lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_employee&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;manager_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;descendant&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;closure&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Subtree rollup (one join, one fixed ancestor).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sales&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;subtree_sales&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;closure&lt;/span&gt; &lt;span class="n"&gt;cl&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;fact_sales&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;employee_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;descendant&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt;  &lt;span class="n"&gt;cl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ancestor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;manager_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;      &lt;span class="c1"&gt;-- add AND cl.depth = 1 for direct reports&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Pattern picker.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fact relates M:N to a dimension&lt;/td&gt;
&lt;td&gt;Bridge between fact's dim and the other dim&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shared measure must sum to the whole&lt;/td&gt;
&lt;td&gt;Bridge + &lt;code&gt;weighting_factor&lt;/code&gt; (sums to 1.0)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One member, many attribute values&lt;/td&gt;
&lt;td&gt;Multi-valued dimension bridge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same value set reused by many members&lt;/td&gt;
&lt;td&gt;Add a group dimension&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Variable-depth tree (org / BOM)&lt;/td&gt;
&lt;td&gt;Closure table + recursive CTE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Count members across a bridge&lt;/td&gt;
&lt;td&gt;&lt;code&gt;COUNT(DISTINCT member_key)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is a bridge table in dimensional modeling?
&lt;/h3&gt;

&lt;p&gt;A bridge table is an associative table placed between two entities that have a &lt;strong&gt;many-to-many&lt;/strong&gt; relationship, holding one row for each valid pair of keys — for example one row per &lt;code&gt;(account_key, customer_key)&lt;/code&gt; when accounts and customers own each other many-to-many. It stores structure, not measures: you join &lt;em&gt;through&lt;/em&gt; it to reach a dimension a fact relates to on both ends. Optional columns on the bridge describe the relationship itself, such as a weighting factor or an effective-date range.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you avoid double-counting through a bridge table?
&lt;/h3&gt;

&lt;p&gt;Double-counting happens because joining a fact through a bridge fans one fact row into several, so a plain &lt;code&gt;SUM&lt;/code&gt; counts a shared value once per related member. Fix it by either aggregating the fact at its own grain before the bridge (for a true grand total), or by adding a &lt;strong&gt;weighting factor&lt;/strong&gt; that sums to 1.0 per parent and summing &lt;code&gt;measure × weight&lt;/code&gt;. For headcounts rather than measures, use &lt;code&gt;COUNT(DISTINCT member_key)&lt;/code&gt; so a member in several groups is counted once.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is a weighting (allocation) factor?
&lt;/h3&gt;

&lt;p&gt;A weighting or allocation factor is a fraction stored on each bridge row that says what share of a parent's measure belongs to that related member. The weights for any one parent must sum to exactly 1.0 — a 50/50 joint account gives each owner 0.5 — so that &lt;code&gt;SUM(measure × weight)&lt;/code&gt; across the many-to-many dimension lands on the true total instead of a duplicated one. It lets you produce an "allocated" report that adds up alongside an unweighted "impact" report that measures reach.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is a multi-valued dimension?
&lt;/h3&gt;

&lt;p&gt;A multi-valued dimension is a dimension member that carries several values of one attribute — a customer in several marketing segments, a visit with several diagnoses, a product with several category tags. You model it with a bridge from the member to a value dimension (&lt;code&gt;dim_customer → bridge_customer_segment → dim_segment&lt;/code&gt;) instead of a comma-separated string or fixed &lt;code&gt;attr_1..attr_n&lt;/code&gt; columns. Filtering on it is safe; counting members across it needs &lt;code&gt;COUNT(DISTINCT)&lt;/code&gt;, and summing a member measure across it needs a weight.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is a closure table and when do I use one?
&lt;/h3&gt;

&lt;p&gt;A closure table (or hierarchy bridge) stores one row for every ancestor-descendant pair in a tree, including each node to itself at depth 0, tagged with the depth between them. Use it for &lt;strong&gt;variable-depth, ragged hierarchies&lt;/strong&gt; — org charts, bills of materials, category trees — where an adjacency list can't be flattened by a fixed number of joins. You build it once with a recursive CTE, then roll a fact up for any subtree with a single join filtered to one ancestor, avoiding an unbounded chain of self-joins.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bridge table vs junk dimension vs many-to-many fact — how do I tell them apart?
&lt;/h3&gt;

&lt;p&gt;A &lt;strong&gt;bridge&lt;/strong&gt; resolves a many-to-many relationship and carries no measures; you join through it. A &lt;strong&gt;junk dimension&lt;/strong&gt; collapses several low-cardinality, unrelated flags (yes/no indicators, small enums) into one dimension to avoid littering the fact with columns — it is not about many-to-many. A &lt;strong&gt;many-to-many fact&lt;/strong&gt; (a factless or transaction fact) records the relationship &lt;em&gt;as an event&lt;/em&gt; and may carry its own additive measures. Reach for a bridge when you need to attach a dimension that relates many-to-many; reach for a factless fact when the relationship itself is the thing you are measuring or counting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every idea above, from the account-customer bridge and the weighting factor that stops double-counting to the multi-valued dimension and the closure-table subtree rollup, maps to a hands-on practice room where you write the SQL against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how do you model a many-to-many between a fact and a dimension without double-counting?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/bridge-table" rel="noopener noreferrer"&gt;Practice bridge-table problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/many-to-many" rel="noopener noreferrer"&gt;Many-to-many modeling drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Special Dimensions: Junk, Degenerate, Role-Playing &amp; Conformed Dimensions</title>
      <dc:creator>Gowtham Potureddi</dc:creator>
      <pubDate>Wed, 23 Sep 2026 13:46:46 +0000</pubDate>
      <link>https://dev.to/gowthampotureddi/special-dimensions-junk-degenerate-role-playing-conformed-dimensions-5eab</link>
      <guid>https://dev.to/gowthampotureddi/special-dimensions-junk-degenerate-role-playing-conformed-dimensions-5eab</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;code&gt;special dimensions&lt;/code&gt;&lt;/strong&gt; are the four named refinements the Kimball method reaches for when a plain star schema starts to sag — when you have a fistful of yes/no flags with nowhere clean to put them, an operational key like &lt;code&gt;order_number&lt;/code&gt; that everybody groups by but that has no attributes to describe, a single calendar that three different date columns all want to join to, and a &lt;code&gt;customer&lt;/code&gt; table that has quietly been redefined in every data mart until no two reports agree. Each smell has a canonical fix, and the fix has a name: the junk dimension, the degenerate dimension, the role-playing dimension, and the conformed dimension.&lt;/p&gt;

&lt;p&gt;None of these is an exception to dimensional modeling; they are dimensional modeling working as intended. The grain of the fact table is still the anchor, surrogate keys still connect facts to dimensions, and a business user can still slice by any attribute. The specials just keep the model honest under pressure — they stop flag columns from sprawling, they stop you from building empty dimension tables around keys, they stop you from copying the date dimension four times, and they stop every mart from inventing its own truth. This guide walks the four patterns an interviewer will actually make you name and defend, and pairs each with a Solution-Tail answer: the DDL, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn54gw6z7hiqdbngui9sn.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn54gw6z7hiqdbngui9sn.jpeg" alt="PipeCode blog header for special dimensions — bold white headline 'Special Dimensions' with subtitle 'Junk · Degenerate · Role-Playing · Conformed' and a stylised star-schema scene with a central fact table linked to four differently-styled dimension cards on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you want &lt;strong&gt;hands-on reps&lt;/strong&gt; immediately after reading, drill the &lt;a href="https://pipecode.ai/explore/practice/topic/junk-dimension" rel="noopener noreferrer"&gt;junk-dimension practice set →&lt;/a&gt;, rehearse operational-key modelling on the &lt;a href="https://pipecode.ai/explore/practice/topic/degenerate-dimension" rel="noopener noreferrer"&gt;degenerate-dimension practice set →&lt;/a&gt;, and wire up multi-role calendars on the &lt;a href="https://pipecode.ai/explore/practice/topic/role-playing-dimension" rel="noopener noreferrer"&gt;role-playing-dimension practice set →&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;On this page&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why special dimensions exist in a star schema&lt;/li&gt;
&lt;li&gt;Junk dimensions — collapse low-cardinality flags into one dim&lt;/li&gt;
&lt;li&gt;Degenerate dimensions — operational keys with no dim table&lt;/li&gt;
&lt;li&gt;Role-playing dimensions — one dim, many roles via views&lt;/li&gt;
&lt;li&gt;Conformed dimensions — shared across facts for drill-across&lt;/li&gt;
&lt;li&gt;Cheat sheet — special-dimension recipes&lt;/li&gt;
&lt;li&gt;Frequently asked questions&lt;/li&gt;
&lt;li&gt;Practice on PipeCode&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Why special dimensions exist in a star schema
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Special dimensions are pattern names for star-schema hygiene problems — learn the smell, then the fix
&lt;/h3&gt;

&lt;p&gt;The one-sentence invariant: &lt;strong&gt;each special dimension is a named answer to a specific way a naive star schema goes wrong, so the skill an interviewer is testing is "hear the smell, name the pattern, write the DDL."&lt;/strong&gt; A star schema is a central fact table (measurements at a declared grain) surrounded by dimension tables (the context you filter and group by), joined through surrogate keys. That works beautifully until the real world hands you attributes that do not fit the tidy "one dimension per business entity" story. The four special dimensions are the tidy answers to those four misfits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The four smells and their fixes.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Flag sprawl → junk dimension.&lt;/strong&gt; You have several low-cardinality, unrelated indicators (&lt;code&gt;is_gift&lt;/code&gt;, &lt;code&gt;is_priority&lt;/code&gt;, &lt;code&gt;order_channel&lt;/code&gt;, &lt;code&gt;payment_type&lt;/code&gt;). Giving each its own dimension is silly (one attribute each), and leaving them as columns on the fact bloats the grain with degenerate text. The junk dimension folds them into one small table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Orphan operational key → degenerate dimension.&lt;/strong&gt; You have a key like &lt;code&gt;order_number&lt;/code&gt; that you must group and count by, but it has no attributes of its own worth a table. You store it &lt;strong&gt;as a column on the fact&lt;/strong&gt; with no dimension behind it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repeated date joins → role-playing dimension.&lt;/strong&gt; &lt;code&gt;order_date&lt;/code&gt;, &lt;code&gt;ship_date&lt;/code&gt;, and &lt;code&gt;due_date&lt;/code&gt; all point at the same calendar. You keep &lt;strong&gt;one&lt;/strong&gt; physical &lt;code&gt;dim_date&lt;/code&gt; and expose it under several view aliases, one per role.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Siloed marts → conformed dimension.&lt;/strong&gt; Sales and returns each built their own &lt;code&gt;customer&lt;/code&gt; table, so cross-process numbers never reconcile. A conformed dimension is one shared, identically-keyed dimension used by every fact, which is what makes drill-across correct.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What stays constant across all four.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Grain first.&lt;/strong&gt; You still declare "one row per ___" for the fact before anything else; the specials never change the grain, they change how context attaches to it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Surrogate keys still rule.&lt;/strong&gt; Junk, role-playing, and conformed dimensions all connect via integer surrogate keys. The degenerate dimension is the deliberate exception — it is a natural key with no surrogate because there is no table to key into.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Additivity is unaffected.&lt;/strong&gt; Special dimensions are about context columns, not measures, so they never touch whether a measure is additive, semi-additive, or non-additive.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What interviewers listen for.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can you &lt;strong&gt;name the pattern from the symptom&lt;/strong&gt; without being handed the term? — the core signal.&lt;/li&gt;
&lt;li&gt;Do you say &lt;strong&gt;"a degenerate dimension has no dimension table"&lt;/strong&gt; and not accidentally build one? — a classic trap.&lt;/li&gt;
&lt;li&gt;Do you explain role-playing as &lt;strong&gt;"one physical table, many views"&lt;/strong&gt; rather than "copy the date dimension"? — senior signal.&lt;/li&gt;
&lt;li&gt;Do you connect conformed dimensions to &lt;strong&gt;drill-across and the bus matrix&lt;/strong&gt;, not just "shared table"? — the whole point of conforming.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Worked example — one order row that needs all four patterns
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; A single fact table can exhibit all four smells at once, which is why they are taught together. Picture &lt;code&gt;fact_sales_line&lt;/code&gt; at the grain of one order line. It carries a handful of flags (junk), an &lt;code&gt;order_number&lt;/code&gt; you group by (degenerate), three dates that all reference the calendar (role-playing), and a &lt;code&gt;customer_key&lt;/code&gt; that must mean the same thing here as it does in &lt;code&gt;fact_returns&lt;/code&gt; (conformed). Seeing all four on one row is the fastest way to internalise that they are complementary, not competing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Given a raw order-line record, label which special-dimension pattern each field belongs to.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;field&lt;/th&gt;
&lt;th&gt;example value&lt;/th&gt;
&lt;th&gt;nature&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;order_number&lt;/td&gt;
&lt;td&gt;SO-1007&lt;/td&gt;
&lt;td&gt;operational key, no attributes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;is_gift / is_priority / channel&lt;/td&gt;
&lt;td&gt;true / false / web&lt;/td&gt;
&lt;td&gt;low-cardinality flags&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;order_date / ship_date / due_date&lt;/td&gt;
&lt;td&gt;2026-03-01 / 03-03 / 03-05&lt;/td&gt;
&lt;td&gt;three refs to one calendar&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;customer_id&lt;/td&gt;
&lt;td&gt;C-42&lt;/td&gt;
&lt;td&gt;shared entity across processes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- The line-item grain, annotated by pattern&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fact_sales_line&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;order_number&lt;/span&gt;   &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;-- degenerate dimension (lives here, no dim table)&lt;/span&gt;
  &lt;span class="n"&gt;junk_key&lt;/span&gt;       &lt;span class="nb"&gt;BIGINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;           &lt;span class="c1"&gt;-- junk dimension (is_gift + is_priority + channel folded)&lt;/span&gt;
  &lt;span class="n"&gt;order_date_key&lt;/span&gt; &lt;span class="nb"&gt;INT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;              &lt;span class="c1"&gt;-- role-playing dim_date (role 1)&lt;/span&gt;
  &lt;span class="n"&gt;ship_date_key&lt;/span&gt;  &lt;span class="nb"&gt;INT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;              &lt;span class="c1"&gt;-- role-playing dim_date (role 2)&lt;/span&gt;
  &lt;span class="n"&gt;due_date_key&lt;/span&gt;   &lt;span class="nb"&gt;INT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;              &lt;span class="c1"&gt;-- role-playing dim_date (role 3)&lt;/span&gt;
  &lt;span class="n"&gt;customer_key&lt;/span&gt;   &lt;span class="nb"&gt;BIGINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;           &lt;span class="c1"&gt;-- conformed dimension (same dim in every fact)&lt;/span&gt;
  &lt;span class="n"&gt;product_key&lt;/span&gt;    &lt;span class="nb"&gt;BIGINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;quantity&lt;/span&gt;       &lt;span class="nb"&gt;INT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;extended_amount&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;order_number&lt;/code&gt; stays as a bare column — it is the degenerate dimension, no &lt;code&gt;dim_order&lt;/code&gt; exists. The three flags collapse into a single &lt;code&gt;junk_key&lt;/code&gt; pointing at &lt;code&gt;dim_order_junk&lt;/code&gt;. The three date columns are three separate foreign keys into the &lt;strong&gt;same&lt;/strong&gt; &lt;code&gt;dim_date&lt;/code&gt;, disambiguated by role. &lt;code&gt;customer_key&lt;/code&gt; points at a &lt;code&gt;dim_customer&lt;/code&gt; that is byte-for-byte the same dimension &lt;code&gt;fact_returns&lt;/code&gt; uses, so the two facts can be joined on it. Every other column is a normal measure or a standard dimension key.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;field&lt;/th&gt;
&lt;th&gt;pattern&lt;/th&gt;
&lt;th&gt;has its own dim table?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;order_number&lt;/td&gt;
&lt;td&gt;degenerate&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;junk_key&lt;/td&gt;
&lt;td&gt;junk&lt;/td&gt;
&lt;td&gt;yes (&lt;code&gt;dim_order_junk&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;order/ship/due_date_key&lt;/td&gt;
&lt;td&gt;role-playing&lt;/td&gt;
&lt;td&gt;yes, one shared &lt;code&gt;dim_date&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;customer_key&lt;/td&gt;
&lt;td&gt;conformed&lt;/td&gt;
&lt;td&gt;yes, shared across facts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; When a field does not fit "one clean dimension per business entity," it is almost always one of these four — walk the list in order (junk, degenerate, role-playing, conformed) and one of them will fit.&lt;/p&gt;

&lt;p&gt;&lt;span&gt;SQL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — dimensional-modeling&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Dimensional-modeling and star-schema design problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/dimensional-modeling" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Schema&lt;/span&gt;
&lt;span&gt;Topic — star-schema&lt;/span&gt;
&lt;strong&gt;Star-schema fact-and-dimension modelling problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/star-schema" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  2. Junk dimensions — collapse low-cardinality flags into one dim
&lt;/h2&gt;
&lt;h3&gt;
  
  
  A junk dimension folds unrelated low-cardinality flags into one table keyed by a single surrogate
&lt;/h3&gt;

&lt;p&gt;The feature that makes a junk dimension worth naming is that it &lt;strong&gt;removes clutter from two places at once&lt;/strong&gt;: it keeps the fact table narrow (one &lt;code&gt;junk_key&lt;/code&gt; instead of five text flags) and it avoids a swarm of one-column dimension tables. "Junk" is Kimball's own slightly jokey term — nothing about the data is low-quality; it is a &lt;em&gt;grab-bag&lt;/em&gt; dimension that gathers the leftover indicators that do not deserve a dimension each.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What belongs in a junk dimension.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Low-cardinality flags.&lt;/strong&gt; Booleans (&lt;code&gt;is_gift&lt;/code&gt;, &lt;code&gt;is_priority&lt;/code&gt;, &lt;code&gt;is_returned&lt;/code&gt;) and short enumerations (&lt;code&gt;order_channel ∈ {web, phone, store}&lt;/code&gt;, &lt;code&gt;payment_type ∈ {card, cash, wire}&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unrelated by nature.&lt;/strong&gt; The flags need not correlate; the junk dimension is a convenience container, not a business entity. It is fine that &lt;code&gt;channel&lt;/code&gt; and &lt;code&gt;payment_type&lt;/code&gt; have nothing to do with each other.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Text you would otherwise dump on the fact.&lt;/strong&gt; Anything that is really a small code or indicator, where storing the raw string on every fact row would bloat the table and slow scans.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How it is populated — observed vs cartesian.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cartesian (full cross-product).&lt;/strong&gt; Enumerate every possible combination up front. For 2 booleans × a 3-value channel × a 3-value payment type that is 2×2×3×3 = 36 rows. Fine when the product is small and bounded.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observed-only (recommended for wide products).&lt;/strong&gt; Insert a combination the first time it actually appears in the data. If only 12 of the 36 combinations ever occur, you carry 12 rows. This is the safer default because the cross-product explodes with each added flag.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The junk_key is a surrogate.&lt;/strong&gt; Each distinct combination gets one integer &lt;code&gt;junk_key&lt;/code&gt;; the fact stores that key, never the raw flags.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;When NOT to junk.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;High-cardinality attributes.&lt;/strong&gt; If a "flag" has thousands of values (a free-text note, a user id), it is not junk — it belongs in its own dimension or is degenerate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attributes analysed independently at scale.&lt;/strong&gt; If the business constantly filters &lt;em&gt;only&lt;/em&gt; by channel across huge scans, a dedicated small dimension (or keeping it standalone) can index better than a combination table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rapidly changing correlated attributes.&lt;/strong&gt; Those point toward a mini-dimension, not a junk dimension.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyoe72ra7cb31nhqqlf1n.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyoe72ra7cb31nhqqlf1n.jpeg" alt="Iconographic junk-dimension diagram — four scattered low-cardinality flag columns on the left collapsing into one dim_order_junk table of observed combinations in the centre, with a single junk_key surrogate replacing all four flags on the fact table on the right." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — building dim_order_junk from observed combinations
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The everyday build is: take the distinct combinations of the flags that actually occur in the source, assign each a surrogate &lt;code&gt;junk_key&lt;/code&gt;, and then look that key up when loading the fact. Below, three flags (&lt;code&gt;is_gift&lt;/code&gt;, &lt;code&gt;is_priority&lt;/code&gt;, &lt;code&gt;channel&lt;/code&gt;) that would otherwise be three columns on every fact row become one small dimension.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Build a junk dimension for &lt;code&gt;is_gift&lt;/code&gt;, &lt;code&gt;is_priority&lt;/code&gt;, and &lt;code&gt;channel&lt;/code&gt;, populated only with combinations that appear in the orders source, and show the resulting table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order&lt;/th&gt;
&lt;th&gt;is_gift&lt;/th&gt;
&lt;th&gt;is_priority&lt;/th&gt;
&lt;th&gt;channel&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;true&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;web&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;web&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;true&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;web&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;true&lt;/td&gt;
&lt;td&gt;phone&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;dim_order_junk&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;junk_key&lt;/span&gt;     &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;GENERATED&lt;/span&gt; &lt;span class="n"&gt;ALWAYS&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;IDENTITY&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;is_gift&lt;/span&gt;      &lt;span class="nb"&gt;BOOLEAN&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;is_priority&lt;/span&gt;  &lt;span class="nb"&gt;BOOLEAN&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;channel&lt;/span&gt;      &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;UNIQUE&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;is_gift&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_priority&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;channel&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;-- one row per distinct combination&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- Observed-only population: insert combinations that actually occur&lt;/span&gt;
&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;dim_order_junk&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;is_gift&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_priority&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;channel&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;is_gift&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_priority&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;channel&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;stg_orders&lt;/span&gt;
&lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;CONFLICT&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;is_gift&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_priority&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;channel&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;DO&lt;/span&gt; &lt;span class="k"&gt;NOTHING&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; The &lt;code&gt;UNIQUE&lt;/code&gt; constraint on the three flag columns is the heart of the pattern — it guarantees exactly one row per distinct combination, so &lt;code&gt;junk_key&lt;/code&gt; is a clean surrogate for "this bundle of flags." The &lt;code&gt;SELECT DISTINCT ... ON CONFLICT DO NOTHING&lt;/code&gt; inserts only combinations present in staging and silently skips ones already loaded, which makes the load idempotent and observed-only. Order 1 and order 3 share the same combination &lt;code&gt;(true,false,web)&lt;/code&gt;, so they collapse to one dimension row.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;junk_key&lt;/th&gt;
&lt;th&gt;is_gift&lt;/th&gt;
&lt;th&gt;is_priority&lt;/th&gt;
&lt;th&gt;channel&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;true&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;web&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;web&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;true&lt;/td&gt;
&lt;td&gt;phone&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; Populate observed-only unless the full cross-product is tiny and bounded — with each extra flag the cartesian product multiplies, and most of those combinations will never occur.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on junk dimensions
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You are loading &lt;code&gt;fact_orders&lt;/code&gt;. The source has &lt;code&gt;is_gift&lt;/code&gt;, &lt;code&gt;is_priority&lt;/code&gt;, and &lt;code&gt;channel&lt;/code&gt; on each order, plus a fully-populated &lt;code&gt;dim_order_junk&lt;/code&gt;. Write the load that replaces those three columns on the fact with a single &lt;code&gt;junk_key&lt;/code&gt;, and explain what happens when a brand-new flag combination shows up.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using a lookup join to resolve junk_key
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Resolve the three flags to one surrogate as the fact loads&lt;/span&gt;
&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;fact_orders&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order_number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;order_date_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;junk_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;order_number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;junk_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                       &lt;span class="c1"&gt;-- the folded flags&lt;/span&gt;
       &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;stg_orders&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_customer&lt;/span&gt;   &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_date&lt;/span&gt;       &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;full_date&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;order_date&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_order_junk&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt;                  &lt;span class="c1"&gt;-- match on the flag bundle&lt;/span&gt;
       &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_gift&lt;/span&gt;     &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_gift&lt;/span&gt;
      &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_priority&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_priority&lt;/span&gt;
      &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;channel&lt;/span&gt;     &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;channel&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;stg_orders row (gift, prio, channel)&lt;/th&gt;
&lt;th&gt;matched junk_key&lt;/th&gt;
&lt;th&gt;fact junk_key stored&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;(true, false, web)&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(false, false, web)&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(false, true, phone)&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(true, true, store)&lt;/td&gt;
&lt;td&gt;none → row dropped by INNER JOIN&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The &lt;code&gt;JOIN dim_order_junk ON (is_gift, is_priority, channel)&lt;/code&gt; translates the three source flags into one &lt;code&gt;junk_key&lt;/code&gt;, which is all the fact stores.&lt;/li&gt;
&lt;li&gt;Existing combinations resolve cleanly; the fact row carries a single integer instead of three columns.&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;new&lt;/strong&gt; combination &lt;code&gt;(true, true, store)&lt;/code&gt; has no matching junk row, so an &lt;code&gt;INNER JOIN&lt;/code&gt; silently drops that fact row — a real bug. The fix is to insert unseen combinations into &lt;code&gt;dim_order_junk&lt;/code&gt; &lt;strong&gt;before&lt;/strong&gt; the fact load (the observed-only &lt;code&gt;INSERT ... DISTINCT&lt;/code&gt; from the worked example) or to use a &lt;code&gt;LEFT JOIN&lt;/code&gt; to a "late-arriving" default key.&lt;/li&gt;
&lt;li&gt;With the dimension populated first, every fact row finds its &lt;code&gt;junk_key&lt;/code&gt; and the flags never touch the fact table.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;fact_orders column&lt;/th&gt;
&lt;th&gt;before pattern&lt;/th&gt;
&lt;th&gt;after pattern&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;flag columns&lt;/td&gt;
&lt;td&gt;is_gift, is_priority, channel&lt;/td&gt;
&lt;td&gt;(removed)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;key column&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;junk_key (one integer)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Junk dimension&lt;/strong&gt;&lt;/strong&gt; — one small table gathers unrelated low-cardinality flags so the fact stays narrow and you avoid a dimension per flag.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Surrogate junk_key&lt;/strong&gt;&lt;/strong&gt; — the fact references a single integer, so a scan reads one column instead of three text/boolean columns per row.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Populate-before-load&lt;/strong&gt;&lt;/strong&gt; — the dimension must contain a combination &lt;em&gt;before&lt;/em&gt; the fact references it, or the inner join drops rows; observed-only insertion up front guarantees coverage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Observed vs cartesian&lt;/strong&gt;&lt;/strong&gt; — inserting only seen combinations keeps the dimension tiny, where the full cross-product would grow multiplicatively with each flag.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the load is O(fact rows) with an index lookup on &lt;code&gt;(is_gift, is_priority, channel)&lt;/code&gt;; the dimension itself is bounded by the number of distinct combinations, typically tens of rows.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;SQL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — junk-dimension&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Junk-dimension and flag-folding problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/junk-dimension" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Modelling&lt;/span&gt;
&lt;span&gt;Topic — low-cardinality&lt;/span&gt;
&lt;strong&gt;Low-cardinality attribute modelling problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/low-cardinality" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  3. Degenerate dimensions — operational keys with no dim table
&lt;/h2&gt;
&lt;h3&gt;
  
  
  A degenerate dimension is an operational key that lives on the fact with no dimension table behind it
&lt;/h3&gt;

&lt;p&gt;The defining property, and the one interviewers trap on: &lt;strong&gt;a degenerate dimension has no dimension table at all — it is a dimension key stored directly in the fact because it has no attributes worth a table.&lt;/strong&gt; The classic examples are transaction identifiers: &lt;code&gt;order_number&lt;/code&gt;, &lt;code&gt;invoice_number&lt;/code&gt;, &lt;code&gt;ticket_id&lt;/code&gt;, &lt;code&gt;shipment_id&lt;/code&gt;, &lt;code&gt;pos_transaction_id&lt;/code&gt;. You genuinely need them — to group line items back into their header, to count distinct orders, to reconstruct a receipt — but there is nothing to &lt;em&gt;describe&lt;/em&gt; about &lt;code&gt;SO-1007&lt;/code&gt; beyond the number itself, so building a &lt;code&gt;dim_order&lt;/code&gt; with a single column would be pure overhead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it is called "degenerate."&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It is a dimension that degenerated to just its key.&lt;/strong&gt; Every attribute that would have lived in a &lt;code&gt;dim_order&lt;/code&gt; has already moved elsewhere: the date went to &lt;code&gt;dim_date&lt;/code&gt;, the customer to &lt;code&gt;dim_customer&lt;/code&gt;, the channel into the junk dimension. What remains is the bare identifier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It stays in the fact, unmodified.&lt;/strong&gt; No surrogate key, no lookup — the natural operational key (&lt;code&gt;order_number&lt;/code&gt;) sits in the fact table as a &lt;code&gt;VARCHAR&lt;/code&gt; (or the source's native type). This is the one place in a star schema where a fact legitimately holds a natural key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is still a dimension conceptually.&lt;/strong&gt; You group by it, filter by it, and count distinct values of it, exactly as you would a normal dimension attribute — it simply has no row storage of its own.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What a degenerate dimension is &lt;em&gt;for&lt;/em&gt;.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Grain marker for header/line models.&lt;/strong&gt; At the order-line grain, &lt;code&gt;order_number&lt;/code&gt; is what ties the lines of one order together. Grouping &lt;code&gt;fact_sales_line&lt;/code&gt; by &lt;code&gt;order_number&lt;/code&gt; reconstructs the order header without a separate header table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distinct counts of transactions.&lt;/strong&gt; &lt;code&gt;COUNT(DISTINCT order_number)&lt;/code&gt; gives you order count from a line-grain fact — a metric you could not compute if the identifier were thrown away.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Traceability back to the source system.&lt;/strong&gt; The operational id lets you join back to the OLTP system for audit or drill-to-detail.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Degenerate vs the alternatives.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Degenerate vs a real dimension.&lt;/strong&gt; If the identifier gains attributes (an order got a status, a priority, a sales rep), those attributes go into a junk dimension or their own dimensions; the id itself stays degenerate. Do not "promote" it to a dimension table just to hang one attribute off it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Degenerate vs junk.&lt;/strong&gt; Junk folds &lt;em&gt;several&lt;/em&gt; low-cardinality flags into a table; degenerate keeps &lt;em&gt;one&lt;/em&gt; high-cardinality operational key with no table. Cardinality and table-existence are the discriminators.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkn3p1uzc7i3wkkhyj0be.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkn3p1uzc7i3wkkhyj0be.jpeg" alt="Iconographic degenerate-dimension diagram — an order_number operational key drawn as a column living directly inside the fact table with no separate dimension table, an X over an empty dim_order box, and header/line rows grouped by the same order_number." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — order_number on a line-grain fact
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The canonical setup is a sales fact at the line-item grain. Each row is one product on one order; the &lt;code&gt;order_number&lt;/code&gt; repeats across the lines that belong to the same order. There is no &lt;code&gt;dim_order&lt;/code&gt; — the number is simply a column — and yet it does all the work of an order dimension.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Model &lt;code&gt;fact_sales_line&lt;/code&gt; so that &lt;code&gt;order_number&lt;/code&gt; is a degenerate dimension, then show how to recover the order-level total from the line grain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_number&lt;/th&gt;
&lt;th&gt;product_key&lt;/th&gt;
&lt;th&gt;quantity&lt;/th&gt;
&lt;th&gt;extended_amount&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SO-1007&lt;/td&gt;
&lt;td&gt;55&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;40.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SO-1007&lt;/td&gt;
&lt;td&gt;81&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;12.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SO-1008&lt;/td&gt;
&lt;td&gt;55&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;20.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fact_sales_line&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;order_number&lt;/span&gt;    &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;-- DEGENERATE DIMENSION: no dim_order table&lt;/span&gt;
  &lt;span class="n"&gt;date_key&lt;/span&gt;        &lt;span class="nb"&gt;INT&lt;/span&gt;     &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;product_key&lt;/span&gt;     &lt;span class="nb"&gt;BIGINT&lt;/span&gt;  &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_product&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;product_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;customer_key&lt;/span&gt;    &lt;span class="nb"&gt;BIGINT&lt;/span&gt;  &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_customer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;quantity&lt;/span&gt;        &lt;span class="nb"&gt;INT&lt;/span&gt;     &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;extended_amount&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- Recover the order header total from the line grain by grouping on the DD&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;order_number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;            &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;line_count&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;extended_amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;order_total&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_sales_line&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;order_number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;order_number&lt;/code&gt; is declared as a plain column with no foreign key — there is nothing to reference because no &lt;code&gt;dim_order&lt;/code&gt; exists. The two &lt;code&gt;SO-1007&lt;/code&gt; rows share the identifier, so &lt;code&gt;GROUP BY order_number&lt;/code&gt; folds them into one header row. &lt;code&gt;SUM(extended_amount)&lt;/code&gt; gives the order total (40.00 + 12.50 = 52.50) and &lt;code&gt;COUNT(*)&lt;/code&gt; gives the line count, both derived purely from the degenerate dimension.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_number&lt;/th&gt;
&lt;th&gt;line_count&lt;/th&gt;
&lt;th&gt;order_total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SO-1007&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;52.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SO-1008&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;20.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If an identifier is needed for grouping or distinct-counting but has no describing attributes, leave it in the fact as a degenerate dimension — building a one-column dimension table for it is a modelling smell, not good hygiene.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on degenerate dimensions
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; A reviewer sees &lt;code&gt;order_number&lt;/code&gt; sitting in your line-grain fact table and says "that natural key belongs in a dimension — build &lt;code&gt;dim_order&lt;/code&gt; and put a surrogate on the fact." Defend the degenerate-dimension design and prove the order-count metric still works.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using the degenerate key directly for order counts
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Order-level KPIs computed straight from the degenerate dimension,&lt;/span&gt;
&lt;span class="c1"&gt;-- no dim_order table, no surrogate key.&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;order_number&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                       &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;extended_amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                               &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;revenue&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;extended_amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;order_number&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;avg_order_value&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_sales_line&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_date&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;month&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;line rows (order_number @ month, amount)&lt;/th&gt;
&lt;th&gt;distinct orders&lt;/th&gt;
&lt;th&gt;revenue&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SO-1007 @ Mar, 40.00&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SO-1007 @ Mar, 12.50&lt;/td&gt;
&lt;td&gt;1 (same order)&lt;/td&gt;
&lt;td&gt;52.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SO-1008 @ Mar, 20.00&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;72.50&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;COUNT(DISTINCT order_number)&lt;/code&gt; collapses the multiple lines of an order into one, giving a correct order count from a line-grain fact — this is exactly the metric a &lt;code&gt;dim_order&lt;/code&gt; would have been built to serve, and the degenerate dimension serves it for free.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;SUM(extended_amount)&lt;/code&gt; aggregates the additive measure across lines for revenue.&lt;/li&gt;
&lt;li&gt;Average order value divides revenue by distinct orders — three KPIs, no order dimension table.&lt;/li&gt;
&lt;li&gt;Building &lt;code&gt;dim_order&lt;/code&gt; with a surrogate would add a join, a load step, and a table that stores nothing but the same number; it buys you nothing because there are no attributes to describe. The degenerate design is simpler &lt;em&gt;and&lt;/em&gt; faster.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;month&lt;/th&gt;
&lt;th&gt;orders&lt;/th&gt;
&lt;th&gt;revenue&lt;/th&gt;
&lt;th&gt;avg_order_value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mar&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;72.50&lt;/td&gt;
&lt;td&gt;36.25&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Degenerate dimension&lt;/strong&gt;&lt;/strong&gt; — the operational key stays in the fact because it has no attributes; there is no dimension table and no surrogate to maintain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Header from line grain&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;GROUP BY order_number&lt;/code&gt; reconstructs the order header from line rows, so you keep the finer grain without losing order-level rollups.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Distinct-count metric&lt;/strong&gt;&lt;/strong&gt; — &lt;code&gt;COUNT(DISTINCT order_number)&lt;/code&gt; is the order-count KPI a &lt;code&gt;dim_order&lt;/code&gt; would exist to provide, delivered directly from the fact.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Avoided join&lt;/strong&gt;&lt;/strong&gt; — no surrogate lookup means one fewer table to load and one fewer join at query time, which matters on billion-row facts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — the distinct-count is O(rows) with a sort/hash on &lt;code&gt;order_number&lt;/code&gt;; a &lt;code&gt;dim_order&lt;/code&gt; would add O(rows) load work and a join for zero descriptive benefit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;SQL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — degenerate-dimension&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Degenerate-dimension and operational-key problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/degenerate-dimension" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Facts&lt;/span&gt;
&lt;span&gt;Topic — fact-table&lt;/span&gt;
&lt;strong&gt;Fact-table grain and measure problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/fact-table" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  4. Role-playing dimensions — one dim, many roles via views
&lt;/h2&gt;
&lt;h3&gt;
  
  
  A role-playing dimension is one physical dimension the fact references several times, each time under a different alias
&lt;/h3&gt;

&lt;p&gt;The whole idea in one line: &lt;strong&gt;when several foreign keys in a fact point at the same dimension for different reasons, you keep one physical dimension and expose it under one view per role — you never copy the table.&lt;/strong&gt; The textbook case is dates: an order has an &lt;code&gt;order_date&lt;/code&gt;, a &lt;code&gt;ship_date&lt;/code&gt;, and a &lt;code&gt;due_date&lt;/code&gt;, and all three are just calendar dates. One &lt;code&gt;dim_date&lt;/code&gt; holds every date once; the fact carries three separate &lt;code&gt;date_key&lt;/code&gt; columns; and three views (&lt;code&gt;dim_order_date&lt;/code&gt;, &lt;code&gt;dim_ship_date&lt;/code&gt;, &lt;code&gt;dim_due_date&lt;/code&gt;) let queries and BI tools join to "the calendar" three times without ambiguity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why views, not copies.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A copy duplicates gigabytes and drifts.&lt;/strong&gt; &lt;code&gt;dim_date&lt;/code&gt; can be large (every day for decades, plus fiscal, holiday, and weekday attributes). Copying it three times triples storage and guarantees the copies fall out of sync when you add an attribute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A view is free and always current.&lt;/strong&gt; &lt;code&gt;CREATE VIEW dim_ship_date AS SELECT * FROM dim_date&lt;/code&gt; is metadata only. Add a column to &lt;code&gt;dim_date&lt;/code&gt; and all three roles inherit it instantly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BI tools need distinct names to disambiguate.&lt;/strong&gt; Most semantic layers cannot join the same physical table to a fact three times and label the joins differently. Distinct view names (&lt;code&gt;dim_ship_date.day_name&lt;/code&gt;) give each role its own unambiguous column set in the model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How the fact is wired.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One foreign key per role.&lt;/strong&gt; The fact has &lt;code&gt;order_date_key&lt;/code&gt;, &lt;code&gt;ship_date_key&lt;/code&gt;, &lt;code&gt;due_date_key&lt;/code&gt; — three integer columns, each a normal FK into &lt;code&gt;dim_date&lt;/code&gt; (or its view alias).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Each role answers different questions.&lt;/strong&gt; Join on &lt;code&gt;order_date_key&lt;/code&gt; to trend bookings, on &lt;code&gt;ship_date_key&lt;/code&gt; for fulfilment, on &lt;code&gt;due_date_key&lt;/code&gt; for SLA breaches — same calendar, three lenses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-role arithmetic is easy.&lt;/strong&gt; Because all three resolve to real dates, &lt;code&gt;datediff(ship_date, order_date)&lt;/code&gt; gives days-to-ship straight from the keys.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Beyond dates.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Role-played &lt;code&gt;dim_employee&lt;/code&gt;.&lt;/strong&gt; A sale can have a &lt;code&gt;sold_by_employee_key&lt;/code&gt; and an &lt;code&gt;approved_by_employee_key&lt;/code&gt; — one &lt;code&gt;dim_employee&lt;/code&gt;, two roles, two views.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Role-played &lt;code&gt;dim_location&lt;/code&gt;.&lt;/strong&gt; Shipments have an &lt;code&gt;origin_location_key&lt;/code&gt; and a &lt;code&gt;destination_location_key&lt;/code&gt; — one &lt;code&gt;dim_location&lt;/code&gt;, two roles.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The rule generalises.&lt;/strong&gt; Any time two FKs from one fact would point at the same dimension, that dimension is role-playing and each role gets a view.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9t7rv0c7b7qwt25ekxji.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9t7rv0c7b7qwt25ekxji.jpeg" alt="Iconographic role-playing-dimension diagram — one physical dim_date table in the centre with three view aliases dim_order_date, dim_ship_date and dim_due_date fanning out, each joined to a separate date_key foreign key on the fact table." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — three date roles over one dim_date
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The standard build is one &lt;code&gt;dim_date&lt;/code&gt;, three FK columns on the fact, and three views. Below, an orders fact records when each order was placed, shipped, and due, all against a single calendar, and the views let you name each role explicitly in a query.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Model &lt;code&gt;fact_orders&lt;/code&gt; with &lt;code&gt;order_date&lt;/code&gt;, &lt;code&gt;ship_date&lt;/code&gt;, and &lt;code&gt;due_date&lt;/code&gt; as role-playing dimensions over one &lt;code&gt;dim_date&lt;/code&gt;, then show days-to-ship per order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_number&lt;/th&gt;
&lt;th&gt;order_date&lt;/th&gt;
&lt;th&gt;ship_date&lt;/th&gt;
&lt;th&gt;due_date&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SO-1007&lt;/td&gt;
&lt;td&gt;2026-03-01&lt;/td&gt;
&lt;td&gt;2026-03-03&lt;/td&gt;
&lt;td&gt;2026-03-05&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SO-1008&lt;/td&gt;
&lt;td&gt;2026-03-02&lt;/td&gt;
&lt;td&gt;2026-03-06&lt;/td&gt;
&lt;td&gt;2026-03-04&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- One physical calendar&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;date_key&lt;/span&gt;  &lt;span class="nb"&gt;INT&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="c1"&gt;-- e.g. 20260301&lt;/span&gt;
  &lt;span class="n"&gt;full_date&lt;/span&gt; &lt;span class="nb"&gt;DATE&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;day_name&lt;/span&gt;  &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;9&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;is_weekend&lt;/span&gt; &lt;span class="nb"&gt;BOOLEAN&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- Three role views over the SAME table (metadata only, no data copied)&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;VIEW&lt;/span&gt; &lt;span class="n"&gt;dim_order_date&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;VIEW&lt;/span&gt; &lt;span class="n"&gt;dim_ship_date&lt;/span&gt;  &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;VIEW&lt;/span&gt; &lt;span class="n"&gt;dim_due_date&lt;/span&gt;   &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fact_orders&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;order_number&lt;/span&gt;   &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;order_date_key&lt;/span&gt; &lt;span class="nb"&gt;INT&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;ship_date_key&lt;/span&gt;  &lt;span class="nb"&gt;INT&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;due_date_key&lt;/span&gt;   &lt;span class="nb"&gt;INT&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- Query each role by its own view alias&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;order_number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="n"&gt;od&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;full_date&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;ordered_on&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="n"&gt;sd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;full_date&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;shipped_on&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;full_date&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;od&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;full_date&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;days_to_ship&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_orders&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_order_date&lt;/span&gt; &lt;span class="n"&gt;od&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;od&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;order_date_key&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_ship_date&lt;/span&gt;  &lt;span class="n"&gt;sd&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;sd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ship_date_key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;dim_date&lt;/code&gt; is created once. The three &lt;code&gt;CREATE VIEW&lt;/code&gt; statements are pure metadata — no rows are copied — and give each role a distinct name so the join is unambiguous. The fact holds three integer FK columns. In the query, joining &lt;code&gt;dim_order_date&lt;/code&gt; and &lt;code&gt;dim_ship_date&lt;/code&gt; (both really &lt;code&gt;dim_date&lt;/code&gt;) resolves the two roles independently, and subtracting the two dates yields days-to-ship per order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order_number&lt;/th&gt;
&lt;th&gt;ordered_on&lt;/th&gt;
&lt;th&gt;shipped_on&lt;/th&gt;
&lt;th&gt;days_to_ship&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SO-1007&lt;/td&gt;
&lt;td&gt;2026-03-01&lt;/td&gt;
&lt;td&gt;2026-03-03&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SO-1008&lt;/td&gt;
&lt;td&gt;2026-03-02&lt;/td&gt;
&lt;td&gt;2026-03-06&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; One physical dimension, one view per role, one FK per role — if you find yourself copying &lt;code&gt;dim_date&lt;/code&gt;, stop and make a view instead.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on role-playing dimensions
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Marketing wants to trend orders by the month they were &lt;strong&gt;placed&lt;/strong&gt;, while operations wants the same fact trended by the month each order was &lt;strong&gt;shipped&lt;/strong&gt;. Using one &lt;code&gt;dim_date&lt;/code&gt;, show both trends and explain why aliasing the dimension is mandatory here.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using two aliased joins to the same dim_date
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Orders placed, by placement month&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;od&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;year&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;od&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;orders_placed&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_orders&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_date&lt;/span&gt; &lt;span class="n"&gt;od&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;od&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;order_date_key&lt;/span&gt;   &lt;span class="c1"&gt;-- role 1&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;od&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;year&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;od&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- Orders shipped, by shipment month, from the SAME fact and SAME dim_date&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;sd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;year&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;orders_shipped&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_orders&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_date&lt;/span&gt; &lt;span class="n"&gt;sd&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;sd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;date_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ship_date_key&lt;/span&gt;    &lt;span class="c1"&gt;-- role 2&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;sd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;year&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;month&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order&lt;/th&gt;
&lt;th&gt;order_date_key (role 1)&lt;/th&gt;
&lt;th&gt;ship_date_key (role 2)&lt;/th&gt;
&lt;th&gt;counted in placed&lt;/th&gt;
&lt;th&gt;counted in shipped&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SO-1007&lt;/td&gt;
&lt;td&gt;20260301 (Mar)&lt;/td&gt;
&lt;td&gt;20260303 (Mar)&lt;/td&gt;
&lt;td&gt;Mar&lt;/td&gt;
&lt;td&gt;Mar&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SO-1008&lt;/td&gt;
&lt;td&gt;20260302 (Mar)&lt;/td&gt;
&lt;td&gt;20260406 (Apr)&lt;/td&gt;
&lt;td&gt;Mar&lt;/td&gt;
&lt;td&gt;Apr&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;The first query joins &lt;code&gt;dim_date&lt;/code&gt; on &lt;code&gt;order_date_key&lt;/code&gt;, so each order is bucketed by the month it was &lt;strong&gt;placed&lt;/strong&gt; — both orders fall in March.&lt;/li&gt;
&lt;li&gt;The second query joins the &lt;strong&gt;same&lt;/strong&gt; &lt;code&gt;dim_date&lt;/code&gt; on &lt;code&gt;ship_date_key&lt;/code&gt;, so orders are bucketed by &lt;strong&gt;ship&lt;/strong&gt; month — SO-1008 shifts to April because it shipped in April.&lt;/li&gt;
&lt;li&gt;Aliasing (&lt;code&gt;od&lt;/code&gt; vs &lt;code&gt;sd&lt;/code&gt;, or the named views) is mandatory: without a distinct alias the optimizer cannot tell which role you mean, and joining one table twice in one query without aliases is a syntax error. The alias &lt;em&gt;is&lt;/em&gt; the role.&lt;/li&gt;
&lt;li&gt;Both trends read from one fact and one physical calendar, so bookings and fulfilment reconcile against the same date attributes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;trend&lt;/th&gt;
&lt;th&gt;Mar&lt;/th&gt;
&lt;th&gt;Apr&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;orders_placed&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;orders_shipped&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Role-playing dimension&lt;/strong&gt;&lt;/strong&gt; — one physical &lt;code&gt;dim_date&lt;/code&gt; answers to several roles, so every date attribute (fiscal period, holiday flag) is defined once and reused everywhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Alias per role&lt;/strong&gt;&lt;/strong&gt; — the table alias or view name is what disambiguates the join; it lets the same dimension attach to the fact multiple times with different meaning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Views over copies&lt;/strong&gt;&lt;/strong&gt; — the role views are metadata, so storage stays flat and adding a &lt;code&gt;dim_date&lt;/code&gt; column propagates to every role instantly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Independent grouping&lt;/strong&gt;&lt;/strong&gt; — grouping by the placement key versus the ship key produces genuinely different trends from the same rows, which is the business value of the pattern.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — each aliased join is O(fact rows) with a key lookup into &lt;code&gt;dim_date&lt;/code&gt;; there is zero extra storage cost because no dimension data is duplicated.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;SQL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — role-playing-dimension&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Role-playing-dimension and multi-date problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/role-playing-dimension" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Schema&lt;/span&gt;
&lt;span&gt;Topic — star-schema&lt;/span&gt;
&lt;strong&gt;Star-schema join and date-dimension problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/star-schema" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  5. Conformed dimensions — shared across facts for drill-across
&lt;/h2&gt;
&lt;h3&gt;
  
  
  A conformed dimension has identical keys and meaning across facts, which is what makes drill-across correct
&lt;/h3&gt;

&lt;p&gt;The invariant to say out loud: &lt;strong&gt;a dimension is conformed when it means exactly the same thing — same keys, same attribute values, same definitions — everywhere it is used, so measures from different fact tables can be combined by that dimension.&lt;/strong&gt; Conformed dimensions are the backbone of Kimball's enterprise bus architecture: instead of each data mart building its own &lt;code&gt;customer&lt;/code&gt; and &lt;code&gt;date&lt;/code&gt; tables (which then never agree), every fact across the organisation shares one &lt;code&gt;dim_customer&lt;/code&gt; and one &lt;code&gt;dim_date&lt;/code&gt;. That shared agreement is precisely what lets you put sales revenue and return volume side by side by customer and trust the result.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What "conformed" actually requires.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Identical surrogate keys.&lt;/strong&gt; &lt;code&gt;customer_key = 42&lt;/code&gt; must be the same customer in &lt;code&gt;fact_sales&lt;/code&gt; and &lt;code&gt;fact_returns&lt;/code&gt;. If the keys differ, the dimension is not conformed and any cross-fact join is meaningless.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Identical attribute meaning.&lt;/strong&gt; &lt;code&gt;customer.segment = 'Enterprise'&lt;/code&gt; must mean the same thing to both processes. Conforming is a governance act, not just a shared table name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same or a strict subset (shrunken).&lt;/strong&gt; A conformed dimension can be the full dimension &lt;em&gt;or&lt;/em&gt; a shrunken rollup of it. A monthly &lt;code&gt;fact_forecast&lt;/code&gt; can conform to &lt;code&gt;dim_date&lt;/code&gt; at the &lt;strong&gt;month&lt;/strong&gt; grain — a "shrunken conformed dimension" whose attributes are a strict subset of the daily one, with keys that roll up cleanly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The bus matrix — how you plan conformance.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Rows are business processes (facts); columns are dimensions.&lt;/strong&gt; Sales, returns, shipments down the side; customer, date, product, store across the top.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A tick means "this fact uses this conformed dimension."&lt;/strong&gt; The matrix is the enterprise data-warehouse plan on one page; shared columns are the conformed dimensions every mart must agree on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conform first, build incrementally.&lt;/strong&gt; You agree the conformed dimensions once, then each fact table can be built independently and still integrate, because they all bolt onto the same dimensions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Drill-across — the payoff.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Never join two facts directly.&lt;/strong&gt; Joining &lt;code&gt;fact_sales&lt;/code&gt; to &lt;code&gt;fact_returns&lt;/code&gt; on &lt;code&gt;customer_key&lt;/code&gt; multiplies rows (a fan trap) and doubles measures. That is the classic mistake.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Split-and-combine.&lt;/strong&gt; Query each fact &lt;strong&gt;separately&lt;/strong&gt;, each aggregated to the &lt;strong&gt;same&lt;/strong&gt; conformed grain (e.g. revenue by customer, returns by customer), then &lt;strong&gt;full-outer-join the two result sets&lt;/strong&gt; on the conformed key. Each measure is computed in isolation, so nothing is double-counted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Common grain is mandatory.&lt;/strong&gt; Drill-across only works where the facts share conformed dimensions at a compatible grain; that shared grain is what the two aggregates line up on.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fov7yk4t5e4fc7nkz4ffl.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fov7yk4t5e4fc7nkz4ffl.jpeg" alt="Iconographic conformed-dimension diagram — one shared dim_customer and dim_date in the centre connected to two separate fact tables fact_sales and fact_returns, with a bus-matrix grid on one side and a drill-across full-outer-join glyph merging the two facts on the common customer key." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4&gt;
  
  
  Worked example — one dim_customer shared by sales and returns
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Detailed explanation.&lt;/strong&gt; The clearest demonstration is two facts and one shared dimension. &lt;code&gt;fact_sales&lt;/code&gt; and &lt;code&gt;fact_returns&lt;/code&gt; both reference the same &lt;code&gt;dim_customer&lt;/code&gt;, so a question like "net revenue = sales minus returns, by customer segment" becomes answerable and trustworthy. The key move is that both facts use the identical &lt;code&gt;customer_key&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; Model &lt;code&gt;fact_sales&lt;/code&gt; and &lt;code&gt;fact_returns&lt;/code&gt; so they both conform to one &lt;code&gt;dim_customer&lt;/code&gt;, then compute sales and returns per customer segment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;dim_customer&lt;/th&gt;
&lt;th&gt;customer_key&lt;/th&gt;
&lt;th&gt;segment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;42&lt;/td&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;43&lt;/td&gt;
&lt;td&gt;SMB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sales rows: (42, 100.00), (43, 30.00). Return rows: (42, 15.00).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;dim_customer&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;     &lt;span class="c1"&gt;-- conformed surrogate: same in every fact&lt;/span&gt;
  &lt;span class="n"&gt;customer_id&lt;/span&gt;  &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;UNIQUE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;segment&lt;/span&gt;      &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fact_sales&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_customer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fact_returns&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_customer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;   &lt;span class="c1"&gt;-- SAME conformed dim&lt;/span&gt;
  &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- Aggregate each fact by the conformed segment attribute&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;sales&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_sales&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;dim_customer&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step explanation.&lt;/strong&gt; &lt;code&gt;dim_customer&lt;/code&gt; is created once with a single &lt;code&gt;customer_key&lt;/code&gt;; both fact tables declare a foreign key to it, so &lt;code&gt;customer_key = 42&lt;/code&gt; is unambiguously the Enterprise customer in both. Aggregating &lt;code&gt;fact_sales&lt;/code&gt; by &lt;code&gt;segment&lt;/code&gt; gives sales per segment; the identical join against &lt;code&gt;fact_returns&lt;/code&gt; gives returns per segment. Because the dimension is shared, the two segment breakdowns line up exactly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;segment&lt;/th&gt;
&lt;th&gt;sales&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;100.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SMB&lt;/td&gt;
&lt;td&gt;30.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb.&lt;/strong&gt; If two facts should ever be compared by a dimension, that dimension must be conformed — build it once, key it once, and point every fact at it.&lt;/p&gt;

&lt;h3&gt;
  
  
  SQL interview question on conformed dimensions
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Question.&lt;/strong&gt; You must report &lt;strong&gt;net revenue (sales − returns) by customer segment&lt;/strong&gt;. A colleague proposes &lt;code&gt;fact_sales JOIN fact_returns ON customer_key&lt;/code&gt;. Explain why that is wrong and write the correct drill-across.&lt;/p&gt;

&lt;h3&gt;
  
  
  Solution Using drill-across with a full outer join on the conformed key
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Code.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Aggregate EACH fact separately to the conformed segment grain,&lt;/span&gt;
&lt;span class="c1"&gt;-- then FULL OUTER JOIN the two results on the conformed key.&lt;/span&gt;
&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="n"&gt;sales&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;sales_amt&lt;/span&gt;
  &lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_sales&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
  &lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_customer&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;
  &lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment&lt;/span&gt;
&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="k"&gt;returns&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;returns_amt&lt;/span&gt;
  &lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_returns&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;
  &lt;span class="k"&gt;JOIN&lt;/span&gt;   &lt;span class="n"&gt;dim_customer&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;
  &lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sa&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;          &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;segment&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sa&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sales_amt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                 &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;sales&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;returns_amt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;               &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;returns&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sa&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sales_amt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;returns_amt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;net_revenue&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;sales&lt;/span&gt; &lt;span class="n"&gt;sa&lt;/span&gt;
&lt;span class="k"&gt;FULL&lt;/span&gt; &lt;span class="k"&gt;OUTER&lt;/span&gt; &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="k"&gt;returns&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sa&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;segment&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step-by-step trace.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;segment&lt;/th&gt;
&lt;th&gt;sales_amt (from sales CTE)&lt;/th&gt;
&lt;th&gt;returns_amt (from returns CTE)&lt;/th&gt;
&lt;th&gt;net&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;100.00&lt;/td&gt;
&lt;td&gt;15.00&lt;/td&gt;
&lt;td&gt;85.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SMB&lt;/td&gt;
&lt;td&gt;30.00&lt;/td&gt;
&lt;td&gt;(none) → 0&lt;/td&gt;
&lt;td&gt;30.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;Directly joining the two facts on &lt;code&gt;customer_key&lt;/code&gt; would pair &lt;strong&gt;every&lt;/strong&gt; sale of a customer with &lt;strong&gt;every&lt;/strong&gt; return of that customer — a fan-out that multiplies rows and inflates both &lt;code&gt;SUM(sales)&lt;/code&gt; and &lt;code&gt;SUM(returns)&lt;/code&gt;. That is the fan trap.&lt;/li&gt;
&lt;li&gt;Drill-across avoids it: each fact is aggregated &lt;strong&gt;on its own&lt;/strong&gt; to the conformed &lt;code&gt;segment&lt;/code&gt; grain first, so each measure is summed exactly once.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;FULL OUTER JOIN&lt;/code&gt; on the conformed attribute stitches the two independent aggregates together and keeps segments that appear in only one fact (SMB has sales but no returns).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;COALESCE(..., 0)&lt;/code&gt; turns the missing side into zero so &lt;code&gt;net_revenue = sales − returns&lt;/code&gt; is correct even when one fact has no rows for a segment.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Output:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;segment&lt;/th&gt;
&lt;th&gt;sales&lt;/th&gt;
&lt;th&gt;returns&lt;/th&gt;
&lt;th&gt;net_revenue&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;100.00&lt;/td&gt;
&lt;td&gt;15.00&lt;/td&gt;
&lt;td&gt;85.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SMB&lt;/td&gt;
&lt;td&gt;30.00&lt;/td&gt;
&lt;td&gt;0.00&lt;/td&gt;
&lt;td&gt;30.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Why this works&lt;/strong&gt; — concept by concept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Conformed dimension&lt;/strong&gt;&lt;/strong&gt; — because &lt;code&gt;dim_customer&lt;/code&gt; and its &lt;code&gt;segment&lt;/code&gt; attribute mean the same thing in both facts, the two aggregates share a common grain that can be lined up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Drill-across&lt;/strong&gt;&lt;/strong&gt; — aggregating each fact separately and joining the results (not the facts) is the only combination method that keeps every measure summed exactly once.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Fan-trap avoidance&lt;/strong&gt;&lt;/strong&gt; — joining facts directly multiplies rows through the shared key; the split-and-combine shape sidesteps it entirely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Full outer join + COALESCE&lt;/strong&gt;&lt;/strong&gt; — preserves segments present in only one process and defaults the missing measure to zero, so net revenue is defined for every segment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/strong&gt; — two independent O(rows) aggregations plus a join on the small segment result set; far cheaper and correct versus the O(sales × returns) blow-up of a direct fact join.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;span&gt;SQL&lt;/span&gt;&lt;br&gt;
&lt;span&gt;Topic — conformed-dimension&lt;/span&gt;&lt;br&gt;
&lt;strong&gt;Conformed-dimension and drill-across problems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/conformed-dimension" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;




&lt;span&gt;Modelling&lt;/span&gt;
&lt;span&gt;Topic — dimensional-modeling&lt;/span&gt;
&lt;strong&gt;Bus-matrix and enterprise-modelling problems&lt;/strong&gt;


&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/dimensional-modeling" rel="noopener noreferrer"&gt;Practice →&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;



&lt;h2&gt;
  
  
  Cheat sheet — special-dimension recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Junk dimension (observed combinations).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;dim_order_junk&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;junk_key&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;GENERATED&lt;/span&gt; &lt;span class="n"&gt;ALWAYS&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;IDENTITY&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;is_gift&lt;/span&gt; &lt;span class="nb"&gt;BOOLEAN&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_priority&lt;/span&gt; &lt;span class="nb"&gt;BOOLEAN&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;channel&lt;/span&gt; &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="k"&gt;UNIQUE&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;is_gift&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_priority&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;channel&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;dim_order_junk&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;is_gift&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_priority&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;channel&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;is_gift&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_priority&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;channel&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;stg_orders&lt;/span&gt;
&lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;CONFLICT&lt;/span&gt; &lt;span class="k"&gt;DO&lt;/span&gt; &lt;span class="k"&gt;NOTHING&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Degenerate dimension (key on the fact, no dim table).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- order_number is a plain column; NO dim_order, NO surrogate&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;order_number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;extended_amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;order_total&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;   &lt;span class="n"&gt;fact_sales_line&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt;  &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;order_number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Role-playing dimension (views over one dim_date).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;VIEW&lt;/span&gt; &lt;span class="n"&gt;dim_order_date&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;VIEW&lt;/span&gt; &lt;span class="n"&gt;dim_ship_date&lt;/span&gt;  &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="c1"&gt;-- fact has order_date_key, ship_date_key, due_date_key -&amp;gt; all -&amp;gt; dim_date&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Conformed dimension (one dim, many facts).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- same customer_key in every fact that references dim_customer&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fact_returns&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;customer_key&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_customer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="nb"&gt;NUMERIC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Drill-across skeleton (never join two facts directly).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;dim&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;m1&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;fact_a&lt;/span&gt; &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="k"&gt;USING&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dim_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;dim&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
     &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;dim&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;SUM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;m2&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;fact_b&lt;/span&gt; &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="k"&gt;USING&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dim_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;dim&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dim&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dim&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;dim&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;m1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;COALESCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;m2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="k"&gt;FULL&lt;/span&gt; &lt;span class="k"&gt;OUTER&lt;/span&gt; &lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dim&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dim&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Pattern picker.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;Special dimension&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Several unrelated low-cardinality flags&lt;/td&gt;
&lt;td&gt;Junk dimension&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operational key (order_number) with no attributes&lt;/td&gt;
&lt;td&gt;Degenerate dimension&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Several FKs to the same dimension (order/ship/due date)&lt;/td&gt;
&lt;td&gt;Role-playing dimension&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One dimension shared across facts / marts&lt;/td&gt;
&lt;td&gt;Conformed dimension&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What are special dimensions in dimensional modeling?
&lt;/h3&gt;

&lt;p&gt;Special dimensions are four named patterns in Kimball dimensional modeling that handle attributes which do not fit the plain "one dimension per business entity" star schema: the &lt;strong&gt;junk dimension&lt;/strong&gt; (folds unrelated low-cardinality flags into one table), the &lt;strong&gt;degenerate dimension&lt;/strong&gt; (an operational key stored on the fact with no dimension table), the &lt;strong&gt;role-playing dimension&lt;/strong&gt; (one physical dimension referenced multiple times under different aliases), and the &lt;strong&gt;conformed dimension&lt;/strong&gt; (one dimension shared identically across many fact tables). They are refinements of the star schema, not exceptions to it — grain and surrogate keys still govern the model.&lt;/p&gt;

&lt;h3&gt;
  
  
  When should I use a junk dimension?
&lt;/h3&gt;

&lt;p&gt;Use a junk dimension when you have several &lt;strong&gt;low-cardinality, unrelated indicators&lt;/strong&gt; — booleans like &lt;code&gt;is_gift&lt;/code&gt; or short enumerations like &lt;code&gt;channel&lt;/code&gt; and &lt;code&gt;payment_type&lt;/code&gt; — that each would make a trivial one-column dimension, and that would bloat the fact if left as text columns. You fold them into a single small table keyed by a surrogate &lt;code&gt;junk_key&lt;/code&gt;, populated with the combinations that actually occur. Do &lt;strong&gt;not&lt;/strong&gt; junk high-cardinality attributes (free text, ids) or attributes the business analyses independently at large scale.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is a degenerate dimension and where is it stored?
&lt;/h3&gt;

&lt;p&gt;A degenerate dimension is a dimension key that has &lt;strong&gt;no attributes and therefore no dimension table&lt;/strong&gt; — it is stored directly as a column in the fact table. The classic examples are transaction identifiers like &lt;code&gt;order_number&lt;/code&gt;, &lt;code&gt;invoice_number&lt;/code&gt;, or &lt;code&gt;ticket_id&lt;/code&gt;. You still group by it (to reconstruct an order header from line rows) and count distinct values of it (to count orders from a line-grain fact), but building a one-column dimension for it would be pure overhead, so the natural key stays in the fact.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do role-playing dimensions work without duplicating data?
&lt;/h3&gt;

&lt;p&gt;You keep &lt;strong&gt;one physical dimension table&lt;/strong&gt; and create a &lt;strong&gt;database view per role&lt;/strong&gt; — for example, one &lt;code&gt;dim_date&lt;/code&gt; with three views &lt;code&gt;dim_order_date&lt;/code&gt;, &lt;code&gt;dim_ship_date&lt;/code&gt;, and &lt;code&gt;dim_due_date&lt;/code&gt;. The fact carries a separate foreign key for each role (&lt;code&gt;order_date_key&lt;/code&gt;, &lt;code&gt;ship_date_key&lt;/code&gt;, &lt;code&gt;due_date_key&lt;/code&gt;), and each query or BI join uses the appropriately named view. Views are metadata only, so no data is copied, and adding a column to &lt;code&gt;dim_date&lt;/code&gt; instantly appears in every role.&lt;/p&gt;

&lt;h3&gt;
  
  
  What makes a dimension conformed?
&lt;/h3&gt;

&lt;p&gt;A dimension is conformed when it means &lt;strong&gt;exactly the same thing everywhere it is used&lt;/strong&gt;: identical surrogate keys, identical attribute values, and identical definitions across every fact table that references it. A shrunken (rollup) version — for example &lt;code&gt;dim_date&lt;/code&gt; at the month grain for a forecast fact — is still conformed if its keys and attributes are a strict subset that rolls up cleanly. Conforming is a governance decision that lets measures from different facts be combined by that dimension.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is drill-across and why does it need conformed dimensions?
&lt;/h3&gt;

&lt;p&gt;Drill-across is the technique for combining measures from &lt;strong&gt;two different fact tables&lt;/strong&gt; by a shared dimension — for example net revenue = sales minus returns by customer segment. You must &lt;strong&gt;not&lt;/strong&gt; join the two facts directly (that fans out rows and double-counts measures); instead you aggregate each fact separately to the same conformed grain, then full-outer-join the two result sets on the conformed key. It only works when the dimension is conformed, because the two aggregates must line up on identical keys and attribute meaning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice on PipeCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/" rel="noopener noreferrer"&gt;Pipecode.ai&lt;/a&gt; is Leetcode for Data Engineering — every special-dimension idea above, from folding flags into a junk dimension to storing a degenerate order key, aliasing one date dimension across roles, and drilling across conformed dimensions, maps to a hands-on practice room where you write the DDL and the query against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you model these three dates against one calendar?" holds up under a senior interviewer's depth probes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pipecode.ai/explore/practice/topic/junk-dimension" rel="noopener noreferrer"&gt;Practice junk-dimension problems now →&lt;/a&gt;&lt;br&gt;
&lt;a href="https://pipecode.ai/explore/practice/topic/conformed-dimension" rel="noopener noreferrer"&gt;Conformed-dimension drills →&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>sql</category>
      <category>interview</category>
      <category>dataengineering</category>
    </item>
  </channel>
</rss>
