<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: dino david</title>
    <description>The latest articles on DEV Community by dino david (@dino_david_c5d8f55f119b7f).</description>
    <link>https://dev.to/dino_david_c5d8f55f119b7f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4128735%2F1e87cc10-230b-4353-986c-19b0e9bff94a.png</url>
      <title>DEV Community: dino david</title>
      <link>https://dev.to/dino_david_c5d8f55f119b7f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dino_david_c5d8f55f119b7f"/>
    <language>en</language>
    <item>
      <title>What a Delta table actually is: Parquet files plus a transaction log</title>
      <dc:creator>dino david</dc:creator>
      <pubDate>Wed, 23 Sep 2026 15:53:35 +0000</pubDate>
      <link>https://dev.to/dino_david_c5d8f55f119b7f/what-a-delta-table-actually-is-parquet-files-plus-a-transaction-log-325b</link>
      <guid>https://dev.to/dino_david_c5d8f55f119b7f/what-a-delta-table-actually-is-parquet-files-plus-a-transaction-log-325b</guid>
      <description>&lt;p&gt;Most Delta Lake explanations start from the feature list - ACID, time travel, schema enforcement - and ask you to take them on trust. Go the other way, start from what is physically on disk, and every one of those features turns out to be the same mechanism wearing a different hat.&lt;/p&gt;

&lt;h2&gt;
  
  
  A table is a directory
&lt;/h2&gt;

&lt;p&gt;There is no database server involved, and no hidden proprietary store. A Delta table is a directory on cloud storage.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data is ordinary Parquet
&lt;/h2&gt;

&lt;p&gt;Every write drops one or more Parquet files into that directory. Parquet is an open columnar format, and crucially, nothing inside a Parquet file knows it belongs to a table. On its own, that directory is just a pile of files.&lt;/p&gt;

&lt;h2&gt;
  
  
  The log is the table
&lt;/h2&gt;

&lt;p&gt;Beside the data files sits a directory named &lt;code&gt;_delta_log&lt;/code&gt;. Every successful write appends one JSON file to it, named by version number and zero-padded to 20 digits - so the first commit is &lt;code&gt;00000000000000000000.json&lt;/code&gt;. One commit file is one version of the table.&lt;/p&gt;

&lt;h2&gt;
  
  
  Inside a commit
&lt;/h2&gt;

&lt;p&gt;Each commit holds a list of actions. The two that do the work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;add&lt;/code&gt;&lt;/strong&gt; - this Parquet file is now part of the table. It carries &lt;code&gt;path&lt;/code&gt;, &lt;code&gt;partitionValues&lt;/code&gt;, &lt;code&gt;size&lt;/code&gt;, &lt;code&gt;modificationTime&lt;/code&gt;, &lt;code&gt;dataChange&lt;/code&gt;, and optionally &lt;code&gt;stats&lt;/code&gt; (per-column statistics).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;remove&lt;/code&gt;&lt;/strong&gt; - this file is no longer part of the table.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Alongside them, &lt;code&gt;metaData&lt;/code&gt; carries the schema and &lt;code&gt;protocol&lt;/code&gt; carries the reader and writer versions a client needs to support. The spec defines others too - &lt;code&gt;commitInfo&lt;/code&gt;, &lt;code&gt;txn&lt;/code&gt;, &lt;code&gt;cdc&lt;/code&gt;, &lt;code&gt;domainMetadata&lt;/code&gt;, &lt;code&gt;sidecar&lt;/code&gt;, &lt;code&gt;checkpointMetadata&lt;/code&gt; - but &lt;code&gt;add&lt;/code&gt; and &lt;code&gt;remove&lt;/code&gt; are the ones that decide what you read.&lt;/p&gt;

&lt;h2&gt;
  
  
  How a read actually works
&lt;/h2&gt;

&lt;p&gt;A reader never just lists the directory and reads what it finds. It reads the log first, replays the &lt;code&gt;add&lt;/code&gt; and &lt;code&gt;remove&lt;/code&gt; actions to work out which Parquet files are current, and only then reads those files.&lt;/p&gt;

&lt;p&gt;That indirection is the whole design. Membership in a table is a log decision, not a filesystem fact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Removed does not mean deleted
&lt;/h2&gt;

&lt;p&gt;A file named in a &lt;code&gt;remove&lt;/code&gt; action is still sitting on storage. It is simply no longer in the current set.&lt;/p&gt;

&lt;p&gt;This is exactly why time travel works: replay the log only up to an older version, and you get that version's file set, still physically present. &lt;code&gt;VACUUM&lt;/code&gt; is what eventually deletes them, and its &lt;strong&gt;default retention threshold is 7 days&lt;/strong&gt; - it removes only files no longer referenced by the table, never files the current version depends on.&lt;/p&gt;

&lt;p&gt;It also explains the warning people run into: vacuum aggressively enough and you delete the files older versions still point at, so time travel stops working for those versions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checkpoints
&lt;/h2&gt;

&lt;p&gt;Replaying thousands of JSON commits on every read would be slow. So Delta periodically writes a checkpoint: a Parquet file in the log directory holding the whole table state at that version. A reader starts from the newest checkpoint and replays only the commits after it. A &lt;code&gt;_last_checkpoint&lt;/code&gt; file points at the most recent one, so the reader does not even need a directory listing to find it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everything else follows
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ACID&lt;/strong&gt; - a commit is one atomic file write. It either lands or it does not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time travel&lt;/strong&gt; - old versions are still described by the log, and their files are still there.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schema enforcement&lt;/strong&gt; - the schema lives in the log, so a write that does not match can be rejected before anything lands.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The one-line answer
&lt;/h2&gt;

&lt;p&gt;A Delta table is Parquet data files plus a transaction log, and every Delta feature is just something that log records.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watch it drawn step by step
&lt;/h2&gt;

&lt;p&gt;The folder opened up one layer at a time, 2:29: &lt;a href="https://www.youtube.com/watch?v=nVNhgjdbYN4" rel="noopener noreferrer"&gt;https://www.youtube.com/watch?v=nVNhgjdbYN4&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Episode 5 of a data engineering interview prep series, in order here: &lt;a href="https://www.youtube.com/watch?v=B5iHmoYgnqY&amp;amp;list=PLDB5WDkDOYF4" rel="noopener noreferrer"&gt;https://www.youtube.com/watch?v=B5iHmoYgnqY&amp;amp;list=PLDB5WDkDOYF4&lt;/a&gt;&lt;/p&gt;

</description>
      <category>deltalake</category>
      <category>databricks</category>
      <category>dataengineering</category>
      <category>bigdata</category>
    </item>
    <item>
      <title>Databricks classic vs serverless compute: the limitation list decides it</title>
      <dc:creator>dino david</dc:creator>
      <pubDate>Mon, 21 Sep 2026 12:49:33 +0000</pubDate>
      <link>https://dev.to/dino_david_c5d8f55f119b7f/databricks-classic-vs-serverless-compute-the-limitation-list-decides-it-527j</link>
      <guid>https://dev.to/dino_david_c5d8f55f119b7f/databricks-classic-vs-serverless-compute-the-limitation-list-decides-it-527j</guid>
      <description>&lt;p&gt;Pick "classic" or "serverless" in Databricks and you are choosing where the machines live. Almost everything else follows from that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same request, twice
&lt;/h2&gt;

&lt;p&gt;You press run on a notebook or a SQL query. Databricks can serve that request two ways.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Classic.&lt;/strong&gt; The control plane asks your cloud provider for machines, and a cluster of VMs is created inside your own account and network. It reads your tables from your own cloud storage. You choose the instance types, you own autoscaling and networking, and the bill has two parts: the cloud provider for the VMs, and Databricks for the DBUs on top.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Serverless.&lt;/strong&gt; The documentation is direct about it: serverless compute resources run in "the serverless compute plane, which is managed by Databricks", and you run workloads "without provisioning any compute resources in your cloud account". Databricks manages, patches and scales that fleet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your data does not move
&lt;/h2&gt;

&lt;p&gt;This is the part that reassures people: in both cases the tables stay in your own cloud storage. It is the compute that changes sides, not the data.&lt;/p&gt;

&lt;h2&gt;
  
  
  The list that actually decides it
&lt;/h2&gt;

&lt;p&gt;Startup time is the headline, but the decision usually comes down to whether your workload uses something serverless does not support. From the limitations page:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;R is not supported.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compute-scoped libraries and compute-scoped init scripts are not supported.&lt;/strong&gt; Notebook-scoped libraries are the alternative.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Only Spark Connect APIs are available&lt;/strong&gt;, and the &lt;strong&gt;Spark RDD APIs are not supported&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Streaming triggers:&lt;/strong&gt; &lt;code&gt;Trigger.AvailableNow()&lt;/code&gt; is supported; &lt;code&gt;Trigger.Continuous(interval)&lt;/code&gt; and &lt;code&gt;Trigger.ProcessingTime(interval)&lt;/code&gt; are not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maximum runtime is 7 days.&lt;/strong&gt; A run that exceeds it is terminated and not retried.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  So which one
&lt;/h2&gt;

&lt;p&gt;Serverless suits SQL warehouses, ad hoc exploration and spiky jobs, where waiting for a cluster is the real cost and nothing on that list applies.&lt;/p&gt;

&lt;p&gt;Classic is the answer when you need something specific: a particular instance type or GPUs, libraries installed at cluster level, R, RDD code, a continuous stream, or your own network setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-line answer
&lt;/h2&gt;

&lt;p&gt;Classic compute runs in your cloud account and you manage the machines; serverless runs in a compute plane managed by Databricks and you manage nothing, while your data stays in your own storage either way.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watch it drawn step by step
&lt;/h2&gt;

&lt;p&gt;Both paths drawn side by side, 2:36: &lt;a href="https://youtu.be/s8452E2Y6WM" rel="noopener noreferrer"&gt;https://youtu.be/s8452E2Y6WM&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Episode 4 of a data engineering interview prep series, in order here: &lt;a href="https://www.youtube.com/watch?v=B5iHmoYgnqY&amp;amp;list=PLDB5WDkDOYF4" rel="noopener noreferrer"&gt;https://www.youtube.com/watch?v=B5iHmoYgnqY&amp;amp;list=PLDB5WDkDOYF4&lt;/a&gt;&lt;/p&gt;

</description>
      <category>databrick</category>
      <category>azure</category>
      <category>dataengineering</category>
      <category>cloud</category>
    </item>
    <item>
      <title>Databricks accounts, workspaces and metastores: which layer owns what</title>
      <dc:creator>dino david</dc:creator>
      <pubDate>Sun, 20 Sep 2026 17:18:29 +0000</pubDate>
      <link>https://dev.to/dino_david_c5d8f55f119b7f/databricks-accounts-workspaces-and-metastores-which-layer-owns-what-4jo3</link>
      <guid>https://dev.to/dino_david_c5d8f55f119b7f/databricks-accounts-workspaces-and-metastores-which-layer-owns-what-4jo3</guid>
      <description>&lt;p&gt;Three words that get used interchangeably until the day you have to design a platform with them. Here is the hierarchy, top down.&lt;/p&gt;

&lt;h2&gt;
  
  
  The account
&lt;/h2&gt;

&lt;p&gt;The account is the container for your whole organisation, and you normally have one per cloud provider. You manage it in the account console, where billing and account admins live. Identities live here too: users, groups and service principals.&lt;/p&gt;

&lt;h2&gt;
  
  
  The workspace
&lt;/h2&gt;

&lt;p&gt;A workspace is a single deployment of the Databricks UI, with its own notebooks, clusters, jobs and dashboards. Most teams run separate workspaces for development, test and production.&lt;/p&gt;

&lt;p&gt;The one-line difference: the account is where you manage identities, billing and metastores for the organisation. A workspace is where people actually do the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Unity Catalog metastore
&lt;/h2&gt;

&lt;p&gt;The metastore is the top-level container for data objects and their permissions, and you create it at account level. The documentation is direct about the count: you must have one metastore for each region in which your organisation operates, and each regional metastore can be linked to any number of workspaces in that region.&lt;/p&gt;

&lt;p&gt;So dev, test and prod in the same region attach to the same metastore, rather than each carrying its own legacy Hive metastore.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three-level namespace
&lt;/h2&gt;

&lt;p&gt;Inside the metastore, objects follow a three-level namespace: a catalog holds schemas, and a schema holds tables, views and volumes. You address a table as &lt;code&gt;catalog.schema.table&lt;/code&gt;, for example &lt;code&gt;prod.sales.orders&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Identity federation
&lt;/h2&gt;

&lt;p&gt;With identity federation you configure users, service principals and groups once in the account console rather than repeating it in every workspace, then assign them to the workspaces that need them. One group, granted permissions on data once, applies everywhere it works. Workspace-local groups still exist in non-federated workspaces, but they cannot be granted Unity Catalog permissions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the data actually sits
&lt;/h2&gt;

&lt;p&gt;Metastore-level managed storage is optional, and managed tables are written there when it is configured. Data you already have in a bucket or container is registered as an external location, backed by a storage credential, and you grant access to that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-line answer
&lt;/h2&gt;

&lt;p&gt;The account holds identities and billing, workspaces hold the tools and the compute, and the metastore holds the data and who can see it, shared across every workspace in the region.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watch it drawn step by step
&lt;/h2&gt;

&lt;p&gt;I drew this as a 2:32 animated diagram: &lt;a href="https://youtu.be/2yy9-BL_RDo" rel="noopener noreferrer"&gt;https://youtu.be/2yy9-BL_RDo&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It is episode 3 of a data engineering interview prep series, in order here: &lt;a href="https://www.youtube.com/watch?v=B5iHmoYgnqY&amp;amp;list=PLDB5WDkDOYF4" rel="noopener noreferrer"&gt;https://www.youtube.com/watch?v=B5iHmoYgnqY&amp;amp;list=PLDB5WDkDOYF4&lt;/a&gt;&lt;/p&gt;

</description>
      <category>databrick</category>
      <category>dataengineering</category>
      <category>azure</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Lakehouse vs Data Warehouse vs Data Lake: The Difference in One Picture</title>
      <dc:creator>dino david</dc:creator>
      <pubDate>Fri, 18 Sep 2026 12:58:26 +0000</pubDate>
      <link>https://dev.to/dino_david_c5d8f55f119b7f/lakehouse-vs-data-warehouse-vs-data-lake-the-difference-in-one-picture-24nl</link>
      <guid>https://dev.to/dino_david_c5d8f55f119b7f/lakehouse-vs-data-warehouse-vs-data-lake-the-difference-in-one-picture-24nl</guid>
      <description>&lt;p&gt;If you work with data, you will be asked this sooner or later: &lt;em&gt;what is a lakehouse, and why not just use a warehouse?&lt;/em&gt; Here is the answer I use, with a short animated diagram at the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data warehouse
&lt;/h2&gt;

&lt;p&gt;A data warehouse (Snowflake, Redshift, BigQuery) stores structured tables in its own optimised format. It is very good at three things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ACID transactions&lt;/li&gt;
&lt;li&gt;enforced schemas and governance&lt;/li&gt;
&lt;li&gt;fast BI queries&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trade-off is cost and reach. Storage is proprietary and expensive at scale, and it was never designed for raw files, images or logs, so machine learning teams struggle to work in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data lake
&lt;/h2&gt;

&lt;p&gt;A data lake is cheap object storage, such as S3 or ADLS, holding open file formats. It takes any data type, costs very little, and Python and ML tools can read it directly.&lt;/p&gt;

&lt;p&gt;But a plain lake has no transactions. Two writers can corrupt a table, nothing enforces a schema, and without a catalog it slowly becomes a data swamp nobody trusts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why most companies ran both
&lt;/h2&gt;

&lt;p&gt;So teams loaded everything into the lake, then copied the modelled part into the warehouse. That means two copies of the data, two pipelines to maintain, and two sets of numbers that drift apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lakehouse
&lt;/h2&gt;

&lt;p&gt;A lakehouse is a data lake with an open table format, such as Delta Lake or Apache Iceberg, on top. That table format gives you warehouse behaviour directly on cheap object storage.&lt;/p&gt;

&lt;p&gt;From the warehouse it takes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ACID transactions on the table files&lt;/li&gt;
&lt;li&gt;schema enforcement and a catalog for governance&lt;/li&gt;
&lt;li&gt;fast BI, because the engine keeps statistics on those same files and skips what a query does not need&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;From the lake it keeps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the same cheap, open storage&lt;/li&gt;
&lt;li&gt;every data type, structured or not&lt;/li&gt;
&lt;li&gt;direct access for Python and ML, with no export step&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The one-line answer
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;One copy of the data, in open formats, serving BI and machine learning at the same time.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Watch it drawn step by step
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/-c09yVpUtAs" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;This is part of a free series that builds Databricks up one layer at a time. Watch it in order from episode 1: &lt;a href="https://www.youtube.com/watch?v=B5iHmoYgnqY&amp;amp;list=PLDB5WDkDOYF4" rel="noopener noreferrer"&gt;https://www.youtube.com/watch?v=B5iHmoYgnqY&amp;amp;list=PLDB5WDkDOYF4&lt;/a&gt;&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>databricks</category>
      <category>deltalake</category>
      <category>bigdata</category>
    </item>
    <item>
      <title>Databricks Architecture Explained: Control Plane, Compute Plane, Delta Lake and Unity Catalog</title>
      <dc:creator>dino david</dc:creator>
      <pubDate>Wed, 16 Sep 2026 21:52:00 +0000</pubDate>
      <link>https://dev.to/dino_david_c5d8f55f119b7f/databricks-architecture-explained-control-plane-compute-plane-delta-lake-and-unity-catalog-1167</link>
      <guid>https://dev.to/dino_david_c5d8f55f119b7f/databricks-architecture-explained-control-plane-compute-plane-delta-lake-and-unity-catalog-1167</guid>
      <description>&lt;p&gt;"Explain the Databricks architecture" is one of the most common data engineering questions. Here is the answer as a picture, built up layer by layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. It starts with your cloud account
&lt;/h2&gt;

&lt;p&gt;Databricks runs on Azure, AWS or Google Cloud, and your cloud account is the foundation everything else sits on.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The control plane
&lt;/h2&gt;

&lt;p&gt;Databricks runs as two halves. The &lt;strong&gt;control plane&lt;/strong&gt; is hosted and managed by Databricks. It is where the web application, the notebooks, the job scheduler and the cluster manager live.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The compute plane
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;compute plane&lt;/strong&gt; (you will also hear "data plane") is where the actual compute runs. In the classic setup it sits inside &lt;strong&gt;your own cloud account&lt;/strong&gt;, in your own virtual network. With serverless compute, it runs in the Databricks account instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Your data lake
&lt;/h2&gt;

&lt;p&gt;At the bottom is your data lake: &lt;strong&gt;ADLS Gen2&lt;/strong&gt; on Azure, &lt;strong&gt;S3&lt;/strong&gt; on AWS. Your raw files live here, and they never leave your account.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Delta Lake
&lt;/h2&gt;

&lt;p&gt;On top of the lake sits &lt;strong&gt;Delta Lake&lt;/strong&gt;, a storage layer that adds ACID transactions, schema enforcement and time travel to plain Parquet files. This is what turns a data lake into a lakehouse.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Bronze, silver, gold
&lt;/h2&gt;

&lt;p&gt;Inside Delta, data is usually organised in three zones:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bronze:&lt;/strong&gt; raw ingested data&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silver:&lt;/strong&gt; cleaned and joined data&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gold:&lt;/strong&gt; business-level aggregates, ready for reporting&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  7. Unity Catalog
&lt;/h2&gt;

&lt;p&gt;Across all of it sits &lt;strong&gt;Unity Catalog&lt;/strong&gt;, the governance layer: permissions, lineage, and a single catalog across every workspace.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. What happens when you run a notebook
&lt;/h2&gt;

&lt;p&gt;The control plane sends the instruction. A cluster spins up in the compute plane, reads from Delta and writes results back. &lt;strong&gt;Only results and metadata go back to the control plane. The data stays put.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  9. The consumers
&lt;/h2&gt;

&lt;p&gt;SQL warehouses for BI, machine learning workloads and streaming pipelines all read from the same gold layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watch it drawn step by step
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/B5iHmoYgnqY" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;This is episode 1 of a free series that builds Databricks up one layer at a time. Watch it in order: &lt;a href="https://www.youtube.com/watch?v=B5iHmoYgnqY&amp;amp;list=PLDB5WDkDOYF4" rel="noopener noreferrer"&gt;https://www.youtube.com/watch?v=B5iHmoYgnqY&amp;amp;list=PLDB5WDkDOYF4&lt;/a&gt;&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>databricks</category>
      <category>azure</category>
      <category>aws</category>
    </item>
  </channel>
</rss>
