Databricks now has a managed Shopify connector in Lakeflow Connect, released in Beta on September 22, 2026. It pulls products, orders, customers, inventory, fulfillments, and discounts from a Shopify store into Unity Catalog tables with incremental syncs, so you no longer need to write and maintain your own API extraction code. It is a strong default for analytics on Shopify data, with a few Beta limitations you should plan around before relying on it in production.
This guide covers how the connector fits into a lakehouse architecture, how to set it up end to end (Shopify app, Unity Catalog connection, and pipeline as code), what to build on top of the raw tables, and where a custom pipeline is still the better choice. Everything here is based on the official Databricks and Shopify documentation as of late September 2026. Because the connector is in Beta, check the linked docs for changes before you build on it.
What does the Lakeflow Connect Shopify connector actually do?
The connector is a managed ingestion pipeline: you point it at a Shopify store, pick the tables you want, and Databricks handles extraction, pagination, rate limiting, incremental cursors, and writing to Delta tables. According to the Shopify connector overview, it supports UI-based and API-based pipeline authoring, Declarative Automation Bundles, incremental ingestion, Unity Catalog governance, orchestration through Databricks Workflows, and SCD Type 2 history tracking.
The pipeline documentation lists 39 source tables in a default source schema. You can ingest individual tables or the whole schema in one declaration.
What it deliberately does not do matters just as much. The limitations page is explicit that the connector ingests raw data without transformations and expects you to handle modeling downstream. It also doesn't cover custom metafields, Shopify reports, or objects outside its 39 prebuilt tables. If your store runs heavily on metafields, that is the first thing to check.
How does it fit into a lakehouse architecture?
The connector replaces the extraction layer, the part most Shopify data teams build by hand with Admin API calls, pagination logic, and retry handling. Everything after it stays the same: bronze tables land in Unity Catalog, you transform them into silver and gold layers, and BI tools, Genie, or ML workloads read from the gold layer.
Two properties of this design are worth calling out.
Credentials live in Unity Catalog, not in code. The Shopify client ID and secret are stored in a Unity Catalog connection object. Pipeline definitions reference the connection by name, so secrets never show up in notebooks, YAML files, or Git history. If you've ever found an access token hardcoded in an old extraction script, this is a real improvement.
Governance applies from the first table. Because the landing tables are Unity Catalog tables, the same access controls, lineage, and auditing you use elsewhere apply to Shopify data immediately. That matters for customer data, which we'll come back to in the limitations section.
What do you need before you start?
The setup splits across Shopify and Databricks, and it's common for a different person to own each side. Check these before starting.
On the Databricks side, the pipeline requirements are:
- A workspace enabled for Unity Catalog, with serverless compute turned on.
-
CREATE CONNECTIONon the metastore to create the connection, orUSE CONNECTIONto use an existing one. -
USE CATALOGon the target catalog, plusUSE SCHEMAandCREATE TABLEon the target schema (orCREATE SCHEMAon the catalog). A workspace admin must enable Lakeflow Connect for Shopify on the Previews page, since the connector is in Beta.
On the Shopify side, per the authentication setup guide:Any Shopify plan works, because the connector reads through the Shopify Admin API, which is available on all plans.
An app created in the Shopify Dev Dashboard, in the same Shopify organization as the store.
The
read_all_ordersscope if you need order history older than 60 days. Shopify grants this scope by approval, so request it early. Without it, you only get the last 60 days of orders.Shopify Payments on the store if you want the
balance_transactionsanddisputestables.
Theread_all_ordersapproval is the step most likely to delay your project, because it depends on Shopify rather than on your team. If historical reporting is the goal, start that request before anything else.
How do you set up the connector step by step?
Setup has three stages: create the Shopify app, create the Unity Catalog connection, and define the ingestion pipeline.
Step 1: Create and install the Shopify app
The connector authenticates with the Shopify OAuth client credentials grant, a machine-to-machine flow with no interactive login. In the Shopify Dev Dashboard:
- Create an app.
- On the app's version, select the read scopes required for the tables you plan to ingest. The Databricks required access scopes list shows what each table needs.
- Install the app on the store.
- Under Settings, copy the Client ID and Client secret.
- Note the store subdomain, which is the
<shop>part of<shop>.myshopify.com. Only grant the scopes for the tables you actually need. A read-only app with a narrow scope list limits the damage if credentials ever leak, and it makes the access easier to justify in a security review.
Step 2: Create the Unity Catalog connection
Only grant the scopes for the tables you actually need. A read-only app with a narrow scope list limits the damage if credentials ever leak, and it makes the access easier to justify in a security review. If an agency manages the store's apps, involve them at this step, since they usually control the Dev Dashboard (for clients of Lucent's Shopify development services, that's our team).
Step 3: Define the pipeline as code
The UI wizard (Data Ingestion → Add data → Shopify) is fine for a first test. For anything you plan to keep, define the pipeline in a Declarative Automation Bundle so it can be reviewed, versioned, and promoted through environments.
The config below ingests three core tables and sets a backfill start date. Per the docs, start_datetime defaults to 365 days before the first sync and only affects that first sync. Later runs resume from the stored cursor.
# resources/shopify_pipeline.yml
resources:
pipelines:
shopify_pipeline:
name: shopify_pipeline
catalog: main
target: shopify_bronze
ingestion_definition:
# References the Unity Catalog connection; no secrets in this file
connection_name: shopify_connection
source_configurations:
- api_source_connector_config:
configs:
# Only applies to the first sync of each incremental table
start_datetime: "2024-01-01T00:00:00+00:00"
objects:
- table:
source_schema: default
source_table: orders
destination_catalog: main
destination_schema: shopify_bronze
destination_table: orders
- table:
source_schema: default
source_table: customers
destination_catalog: main
destination_schema: shopify_bronze
destination_table: customers
- table:
source_schema: default
source_table: products
destination_catalog: main
destination_schema: shopify_bronze
destination_table: products
Add a job file to control the refresh schedule. This one runs daily at midnight UTC:
# resources/shopify_job.yml
resources:
jobs:
shopify_job:
name: shopify_job
schedule:
quartz_cron_expression: "0 0 0 * * ?"
timezone_id: UTC
tasks:
- task_key: shopify_ingestion
pipeline_task:
pipeline_id: ${resources.pipelines.shopify_pipeline.id}
Deploy with databricks bundle deploy, starting with a development target. Run the first backfill against a dev or staging catalog before you point anything at production. A year of order history on a large store is a big first sync, and you want to confirm the table shapes before dashboards depend on them.
A note on the backfill date: if the app doesn't have read_all_orders, a start_datetime older than 60 days won't give you older orders. The date and the scope need to agree.
What should you build on top of the raw tables?
The connector lands raw data, so the next job is shaping it into something analysts and dashboards can use. The limitations page recommends Spark Declarative Pipelines for downstream transformations, and a materialized view is a simple starting point for a daily sales table.
The example below is illustrative. Check the connector reference for exact column names and types before you use it, since Beta schemas can change.
-- Gold-layer daily sales view (column names are placeholders:
-- verify them against the connector's table reference)
CREATE OR REPLACE MATERIALIZED VIEW main.shopify_gold.daily_sales AS
SELECT
DATE(o.created_at) AS order_date,
COUNT(DISTINCT o.id) AS orders,
SUM(CAST(o.total_price AS DECIMAL(18,2))) AS gross_sales,
COUNT(DISTINCT o.customer_id) AS unique_customers
FROM main.shopify_bronze.orders AS o
WHERE o.cancelled_at IS NULL
GROUP BY DATE(o.created_at);
Two modeling decisions are worth making early.
Gross versus net sales. The view above reports gross sales. Net sales needs refunds and discounts subtracted, and those come from their own tables. Settle on one definition and document it, because a finance team and a marketing team pulling "revenue" from different tables is the most common source of dashboard arguments.
Customer history with SCD Type 2. With SCD Type 2 turned on, you keep every version of a customer record instead of only the latest. That lets you answer questions like which segment a customer belonged to when they placed an order. Keep in mind the Beta caveat below: tables with semi-structured columns can't use it.
What are the Beta limitations you should plan around?
Most of these are reasonable for a Beta connector, but a few can surprise you in production. The table below collects them from the overview, limitations, and troubleshooting pages.
| Limitation | What it means in practice |
|---|---|
| 39 prebuilt tables only; no custom metafields or reports | Stores that model key data in metafields need a separate path for those fields |
locations, shop, collects, disputes, countries, collection_product, and inventory_level aren't incremental |
Every run re-ingests these tables in full, so watch runtime and cost on frequent schedules |
| No SCD Type 2 on tables with VARIANT columns | Plan history tracking table by table, not as a blanket setting |
| Schema evolution handles new and deleted columns, but not type changes | A type change upstream needs manual handling |
| Column renames require a full refresh | Budget time for a re-sync when Shopify renames a field |
| Adding a column later doesn't backfill it | Run a full refresh on that table if you need its history |
| Renaming a destination table makes the pipeline API-only | Decide naming up front if your team relies on the UI |
| No API-based row filtering | You ingest whole tables and filter downstream |
| Alerts on scheduled pipelines fire on the next update, not immediately | Don't use pipeline alerts as real-time incident monitoring |
Two more points deserve attention.
Protected customer data. The setup guide notes that some customer and order fields are protected customer data, and Shopify only returns them from non-development stores once the app meets its protected customer data requirements. If email or address fields come back empty in production but not in development, this is the likely cause. Once those fields do arrive, treat them as PII from day one. Unity Catalog's attribute-based access control can mask columns such as email and phone, and Databricks added metastore-level ABAC policies in Beta this month, so one masking rule can cover every catalog. Doing this before analysts get access is far easier than retrofitting it for GDPR requests later.
Rate limits. Shopify rate limits its GraphQL Admin API by query cost, with limits that depend on your plan. According to the troubleshooting guide, the connector waits and retries automatically when it's throttled. That's convenient, but other apps on the same store share that budget. A heavy backfill during peak trading hours can slow down other integrations, so schedule the first sync for a quiet period.
When should you still build a custom Shopify pipeline?
The managed connector should be your default for analytics ingestion. It removes a lot of code you'd otherwise have to maintain, and it gets governance right from the start. But it solves one problem well, and some requirements sit outside it.
| Requirement | Managed connector | Custom pipeline |
|---|---|---|
| Scheduled analytics on core Shopify objects | Best fit | Unnecessary effort |
| Metafield-heavy data models | Not covered | Needed for those fields |
| Reacting to events within seconds (order confirmations, fraud checks) | Scheduled syncs | Webhooks with a queue |
| Writing data back to Shopify | Read-only | Required |
| Logic across many stores, such as shared catalogs | One connection per store | Often simpler to own |
In practice, many teams will end up with a hybrid. The connector handles bulk analytics ingestion, and a small custom service handles metafields or real-time event reactions. If you've built webhook pipelines before, the reliability checks in our Shopify webhook debugging checklist still apply to that second half. The only difference is that your service no longer has to double as your analytics pipeline.
Where the boundary sits depends on the store. It comes down to how much of the data model lives in metafields, how fresh the data needs to be, and how many stores are involved. That trade-off analysis is most of what Lucent works through with commerce clients as a Databricks-focused data engineering company, and it's worth doing on paper before writing any pipeline code, whoever builds it.
How should you roll this out safely?
Because the connector is in Beta, treat the rollout as a staged migration rather than a switch.
- Enable the preview in a non-production workspace first. Confirm the tables, scopes, and backfill behavior before production depends on them.
-
Request
read_all_ordersearly if historical reporting matters. - Run both pipelines in parallel if you're replacing a custom one. Compare row counts and key totals for a week or two before retiring the old pipeline, and keep it available as a rollback until the numbers match.
- Apply PII masking before granting analyst access, not after.
- Pin the pipeline in a bundle so every change goes through code review. The Shopify connector removes the least interesting part of commerce analytics: extraction code that breaks whenever an API changes. What's left is the valuable work of modeling revenue correctly, protecting customer data, and deciding what really needs to be real time.
Are you planning to move an existing Shopify pipeline onto the managed connector, or keep a hybrid setup? I'd especially like to hear from anyone whose store relies heavily on metafields, since that's the gap I expect most teams to hit first.



Top comments (0)