DEV Community

Cover image for dbt Is Leaving the Warehouse. Are Your Pipelines Ready?
Andrew Tan
Andrew Tan

Posted on Originally published at layline.io

dbt Is Leaving the Warehouse. Are Your Pipelines Ready?

dbt built the modern analytics workflow. Now the transformation layer is moving upstream — and the gap between SQL models and runtime execution is becoming the bottleneck nobody planned for.

The scheduled run that used to be enough

For most of the last decade, the analytics workflow looked like this:

Extract data from your sources. Load it into a warehouse. Write SQL models in dbt. Schedule them to run every few hours. Build dashboards on top. When the business asks why the numbers look off, you check the freshness badge and discover the job failed six hours ago.

It wasn't perfect, but it was predictable. The contract between analytics and engineering was clear: models are built in the warehouse, on a schedule, in SQL. dbt made that contract elegant. Version-controlled models, automated testing, dependency graphs, documentation generated from code. For batch analytics, it was a genuine leap forward.

Then the business started asking for fresher data. Not next-day fresh. Not hourly fresh. Event-driven fresh. The kind of fresh where a customer updates their profile and the recommendation engine knows about it before they click away.

And suddenly, the scheduled run isn't enough.


Why dbt's success created a runtime gap

dbt's core insight was separating model logic from execution infrastructure. You write the SQL. dbt handles the DAG, the tests, the documentation, the materialization strategy. The actual compute runs wherever you point it — Snowflake, BigQuery, Redshift, Databricks. dbt doesn't care. It's a compiler and orchestrator for SQL models, not a runtime.

That separation was genius for batch workloads. It meant analytics engineers could own modeling logic without managing clusters, partitioning strategies, or memory tuning. The warehouse handled all of that. The analytics engineer focused on semantics: what does this model mean, how is it tested, who depends on it.

But that separation assumes something important: that the execution layer can handle the workload you're throwing at it. And for scheduled batch jobs against a modern warehouse, that's true. For real-time processing against streaming data, it's not.

The gap isn't in the SQL. dbt's SQL is still SQL. The gap is in everything that happens between the event and the model result:

  • Freshness SLA. A dbt model running every fifteen minutes is still fifteen minutes behind. In a streaming context, fifteen minutes is a batch job wearing a smaller costume.
  • Out-of-order events. Streamed data arrives late, duplicated, or out of sequence. A SQL model that assumes ordered input produces wrong results without warning.
  • Stateful operations. Window functions, sessionization, deduplication — these need to maintain state across events. dbt's materialization strategies (table, incremental, ephemeral) weren't designed for event-time state management.
  • Operational joins. Joining a stream of orders to a slowly changing dimension table in real time is a different problem than joining two warehouse tables on customer_id. The dimension might change mid-query. The stream doesn't wait.

These aren't edge cases. They're the defining characteristics of stream processing. And dbt, by design, delegates all of them to the execution layer.


What the vendors are actually selling

If you've been to a data conference in the last year, you've seen the marketing shift. Confluent is pushing dbt-native modeling on Flink. Fivetran launched dbt Wizard with AI-assisted model generation. Snowflake announced Dynamic Tables, BigQuery has Materialized Views, Databricks is doubling down on Delta Live Tables.

The message is consistent: you can keep your dbt workflow and just... make it real-time. The SQL stays the same. The DAG stays the same. The only thing that changes is the speed.

This is approximately half true.

Yes, you can run SQL on streaming data. Flink SQL, Spark Structured Streaming, and a growing list of stream processors support SQL syntax that looks familiar. Yes, you can version-control those SQL files and build DAGs. Some tools even support dbt-style testing and documentation.

What they don't tell you is that the other half of the problem — the runtime half — doesn't go away. It just changes shape.

When you move from scheduled batch to continuous streaming, you inherit a new set of concerns that dbt never had to solve:

Batch assumption Streaming reality
Data is complete when the job starts Data is never complete; late arrivals are normal
Failures are caught at the end of the run Failures must be detected and handled per-event
Schema changes happen between runs Schema changes happen mid-stream
Reprocessing means rerunning a job Reprocessing means rewinding a stream and replaying
Cost is proportional to data volume Cost is proportional to infrastructure uptime

The SQL might look the same. But the system running it is solving an entirely different set of problems.


The ownership question nobody wants to answer

Here's where it gets organizational. dbt created a clean boundary: analytics engineers own the models, platform engineers own the warehouse. When a model fails, it's usually a SQL problem or a data quality problem. When the warehouse is slow, it's a platform problem. The separation of concerns mapped neatly to a separation of teams.

Streaming breaks that boundary.

When a real-time pipeline fails, is it a SQL problem or an infrastructure problem? If events are arriving out of order, does the analytics engineer rewrite the window logic, or does the platform engineer reconfigure the stream processor's watermark settings? If a late-arriving event corrupts a materialized view, who fixes it — the person who wrote the SQL or the person who manages the checkpointing?

I've seen teams handle this three ways, and only one of them works:

Option 1: Analytics engineers learn stream processing. They become fluent in checkpointing, backpressure, event time vs. processing time, and exactly-once semantics. This works for small teams with senior people. It doesn't scale.

Option 2: Platform engineers own everything downstream of Kafka. The analytics engineers write SQL specs, and the platform team implements them in Flink or Spark. This preserves the separation but creates a translation layer. SQL specs are ambiguous about event-time behavior. The platform team makes assumptions. Those assumptions become bugs six months later.

Option 3: Separate the modeling logic from runtime responsibility. Analytics engineers own what the model means — the business logic, the tests, the semantics. Platform engineers own how it runs — the execution engine, the state backend, the failure recovery. Both sides agree on a contract: the model expects ordered input within a bounded lateness, and the runtime guarantees that contract or surfaces a clear error.

This third option is harder to set up. It requires both teams to agree on interfaces and failure modes upfront. But it's the only one that scales without turning your analytics engineers into distributed systems specialists or your platform team into mind readers.


What "real-time dbt" actually requires

If your team is serious about moving dbt-style models into real-time pipelines, here is what you actually need — not what's in the marketing deck:

A runtime that understands event time. Not just processing time. Event time. The difference between "when did this arrive" and "when did this happen" is the difference between correct results and subtly wrong results that look fine on a dashboard.

Explicit state management. Windowing, deduplication, and sessionization all require state. That state needs to be checkpointed, recoverable, and queryable for debugging. If you can't inspect what the system thought at 2:47 PM last Tuesday, you can't debug a production incident.

Schema evolution with teeth. Adding a column is easy. Handling a type change, a rename, or a semantic shift in a column's meaning is hard. Your pipeline needs to detect these changes, decide whether they're safe, and either adapt or stop with a clear error.

Replay and backfill as first-class operations. In batch, reprocessing is rerunning a job. In streaming, it's rewinding a log and replaying events through the same logic. If your real-time pipeline can't replay exactly, you can't recover from bugs, you can't test changes against historical data, and you can't prove compliance.

Cost visibility by workload. Batch costs are easy to reason about: this job processed this much data and took this long. Streaming costs are continuous: the job is always running, always consuming resources, and the cost doesn't correlate cleanly with business output. You need telemetry that connects infrastructure spend to pipeline behavior.

None of these are SQL features. They're runtime features. And they're the difference between a demo that runs for ten minutes and a system that runs for ten months.

A team of engineers standing at the boundary between a calm warehouse room and a chaotic streaming pipeline room, with one person trying to hold a SQL model that is stretching and deforming between the two spaces

The honest pitch

dbt changed how analytics teams work. It brought software engineering practices — version control, testing, documentation — to a discipline that badly needed them. That contribution is real and lasting.

But dbt was built for a world where data moves in batches, warehouses are the center of gravity, and "fresh" means "updated this hour." The industry is moving toward a world where data moves continuously, models are applied in flight, and "fresh" means "updated this millisecond."

That doesn't make dbt obsolete. It makes dbt incomplete for a growing set of use cases.

The vendors selling "real-time dbt" are responding to a real demand. But what they're selling is usually a SQL interface on a stream processor, not a solution to the runtime problems that stream processing introduces. The SQL is the easy part. The state management, failure recovery, schema evolution, and operational observability are the hard parts. And those hard parts don't disappear just because the SQL looks familiar.

If you're evaluating one of these tools, don't ask "Can it run my dbt models faster?" Ask "What happens when a node restarts mid-window?" "How do I replay last Tuesday's data through a model I changed yesterday?" "What does the system do when schema changes at 2 AM?"

The answers to those questions will tell you whether you're buying a faster batch tool or a real streaming runtime.


Where we fit in

At layline.io, we built a runtime that handles both batch and streaming in the same pipeline. Not two separate systems with a SQL veneer over each. One system where the same team can build scheduled warehouse loads and real-time event processing without switching tools, contexts, or mental models.

The analytics engineers keep ownership of model semantics. The platform team keeps ownership of execution infrastructure. But both sides work in the same environment, with the same observability, and the same guarantees around replay, state recovery, and schema evolution.

dbt taught the industry that modeling logic deserves rigor. We're building on that idea — and adding the runtime rigor that real-time data demands.

Top comments (0)