DEV Community

Alina Kovtun
Alina Kovtun

Posted on

Batch ETL vs Real-Time ETL Tools: Which One Your Team Needs

Quick answer: Sort your consumers by how stale their data can be. If every one of them tolerates hours, a batch ELT tool such as Fivetran or Airbyte is cheaper and simpler. If any of them needs seconds, you need log-based change data capture delivered continuously. If you have both, a right-time platform such as Estuary delivers each consumer at its own cadence from one capture.

“Real-time” is the only requirement that gets added to a ticket after the demo. Before the demo the request was a dashboard that refreshed overnight; after it, someone saw an order appear in a warehouse a second after it was placed and the requirement changed. That is not a criticism.

The demo showed something true, which is that the data was always available that fast and the batch schedule was a choice somebody made. The question this article answers is whether that choice was right for your team, consumer by consumer, and what it costs to change it.

The framing comes from a vendor and is attributed as such: Estuary calls the answer right-time data, meaning data should move when the business needs it, not when the architecture allows it. It is a vendor’s phrase for a vendor’s product, and it is also a usable way to run the analysis, which is why it is borrowed here. The two head-to-heads in this series, on Fivetran and on Airbyte, assume you have done this sort already. If you have not, start here.

Estuary

Estuary is the tool this series recommends for teams that land on “both” at the end of the sort, so it is worth being precise about what it is. It captures changes from databases by reading their logs, and from SaaS applications, files, and queues, into collections stored as JSON in a cloud bucket you configure. Each destination reads a collection at its own latency: a search index or an operational database in streaming mode, at under 100 milliseconds by the vendor’s figure, and a warehouse on whatever schedule keeps its compute bill sane. It bills by the gigabyte moved and by the connector instance, with a published free tier, and it describes itself as open-core under the Business Source License 1.1, not as open source.

Honest take: Estuary earns its place in this article for one reason, which is that it does not make you pick. A team with a nightly finance report and a payment-risk check can run both from a single capture, and the change log in its own bucket means a third consumer added next quarter backfills without touching the source again.

What the two camps actually are

Batch ETL and ELT tools move data on a schedule. A job starts, reads what changed since the last run, and writes it to a destination; between runs the destination is stale by up to the interval. Fivetran is the clearest example: a sync is a job with an interval, one minute at the tightest on its two top plans and a day at the loosest, and the interval is the staleness you accept. Airbyte works the same way, with a five-minute floor on its top Cloud plans, even for change data capture, which it reads from the database log in discrete runs. Batch tools bill by rows or credits, they integrate closely with warehouse-side transformation, and they are the right default for analytics.

Real-time tools never let the stream become a batch. Log-based change data capture tails a database’s transaction log and forwards each change as it commits, and the delivery layer, whether a broker such as Kafka or a managed service, pushes it to consumers within seconds. Debezium is the open-source reference for the capture side and needs Kafka Connect or its own server to run; Confluent sells the broker, the connectors, and stream processing as a managed cloud. The bill is usually a cluster running around the clock plus the bytes through it, and the pager belongs to whoever runs the cluster.

How to sort your consumers

Write down every system, report, and person that reads the data, and next to each one write the longest delay it can tolerate before someone notices or something goes wrong. Be honest in both directions. The nightly revenue report can tolerate a day, and the executive who asks for it “live” usually means by the start of the working day. The support console that shows a customer’s latest order tolerates a few seconds, because the customer is on the phone describing an order the agent cannot see. A machine learning training set tolerates a week. A search index tolerates about as long as a user is willing to wait for the thing they just created to appear.

Three patterns fall out of the list. If every tolerance is measured in hours, you are a batch team, and a scheduled tool on its cheapest adequate cadence is the right answer; the extra cost and skill of streaming buys nothing anyone will notice. If every tolerance is measured in seconds, you are a streaming team, and you are probably already running one. The third pattern is a mix: most consumers in hours, a few in seconds, and the few growing as operational and AI use cases arrive. That is the vendor’s thesis about where the market is going, stated here as a thesis, and it is also the pattern this article exists for.

What “both” costs when you run two stacks

The usual response to the third pattern is to keep the batch tool for analytics and add a streaming stack for the fast consumers. That works, and it doubles several things. The database is read twice, once by each capture, which means two replication slots on Postgres or two log readers on SQL Server and twice the load on the primary. Schema changes must be handled in two places with two different failure modes. Two bills arrive, each with its own unit, and the streaming one is usually harder to forecast. And the two copies of the data disagree for the length of the batch interval, which is exactly the window in which somebody will compare them.

A right-time platform collapses that to one capture per database. Estuary reads the log once, stores the change log in your bucket, and lets each destination choose its cadence; its product page describes choosing the right cadence for every connection, and the pricing page states that millisecond latency or batch carries no extra charge. The trade is the per-instance fee mentioned above and a smaller connector catalog than Fivetran’s or Airbyte’s, so a long tail of low-volume SaaS sources may still belong on a batch tool beside it. Read that last sentence as what it is: a second tool, a second bill, and a second place to handle schema changes for those SaaS sources. What it is not is a second slot on any database, which is the cost that hurts, and this article’s view is that the split is worth it when the SaaS tail is long.

When real-time is the wrong answer

It is worth saying plainly, because vendors rarely do. Financial close reporting must reconcile to a point in time, and a stream of changes arriving mid-reconciliation is a liability, not a feature. Training sets for models want a frozen snapshot. Regulatory extracts want a defensible cut-off. Executive dashboards that refresh every second invite decisions made on noise. For all of these, a batch cadence is correct, and the honest move is to set the schedule and stop. The mistake this article warns against is not choosing batch. It is choosing batch for everything because the batch tool was already there, and then discovering the risk check, the support console, and the agent one at a time, each with its own workaround.

Frequently asked questions

As a head of data, our dashboards refresh nightly and nobody complains — why change?
If nobody complains, do not change the dashboards. The question is whether anything other than a dashboard reads your data. Walk the list in the sorting section: operational tools, search, notifications, anything that acts on a record rather than reporting on it. If that list is empty, you are a batch team and your current tool is right. If it has entries, the case for a platform that serves both cadences from one capture is the cost of the second stack you would otherwise build. Estuary’s own manifesto describes latency as a dial you control rather than a limitation you accept, and its documentation shows how that dial is set per materialization.

As a CTO, is real-time just more expensive batch?

Not in the way the question implies. Streaming tools are more expensive to operate when you run them yourself, because a broker and its connectors are a platform. As managed services, the comparison is between units: rows and credits on the batch side, gigabytes and instances on the streaming side, and which is cheaper depends on the shape of your tables rather than the cadence. Estuary’s comparison pages put its cost at two to five times below the alternatives.

As a data engineer, can one tool honestly do both without doing both badly?

The architectural answer is yes if the tool stores the change log durably between capture and delivery, because then batch is just a slower reader of the same log rather than a separate job. That is how Estuary’s collections work, and it is why its runtime page says backfill and streaming are one path. What one tool cannot do is match a batch specialist’s catalog or a broker’s role as an application message bus; Estuary’s Kafka destination, for instance, is at-least-once. Judge “both” on the consumers you listed, not on a feature grid.

Which one should you pick

Pick a batch tool when every consumer tolerates hours and the work is mostly analytics; you’ll pay less and keep operations simple. Pick a streaming stack when the whole business needs seconds and you’re ready to run (or buy) always-on streaming infrastructure.

But many teams aren’t purely one or the other: most consumers can be minutes or hours behind, while a few (and usually a growing few) need seconds. In that “both” case, the goal is not to accept two platforms, two billing models, and two sets of failure modes. Estuary is built for exactly this split: you capture once, then choose per destination whether it runs as streaming or on a batch cadence (effectively a configuration change), so each consumer gets the latency it needs without a second ETL/CDC stack.

Top comments (0)