DEV Community

Mohammed Arshad Ansari
Mohammed Arshad Ansari

Posted on Originally published at hikmahtechnologies.com

DuckDB vs ClickHouse: which one, and when you need both

DuckDB and ClickHouse get compared because they sit near each other in every analytical benchmark: both are columnar, both are vectorised, both are very fast. The benchmark framing hides the difference that actually decides it. DuckDB is a library you put inside your program. ClickHouse is a server your programs connect to. Almost every practical difference follows from that.

I run both. ClickHouse is the analytical store behind Ansaar, feeding a live API. DuckDB is what I reach for in batch jobs and local work, and it is the subject of my book. So this is the comparison from operating them, not from a benchmark table.

The short answer

Pick DuckDB when the work happens inside one program — a transformation job, a service that owns its own queries, a notebook, a CI run — and the data arrives in batches. There is nothing to operate, and it is extraordinarily fast for that one program.

Pick ClickHouse when many clients query the same data at the same time, the data keeps arriving, and people expect the numbers to be seconds old. That is a server's job, and ClickHouse is one of the best servers for it.

If you have both shapes, run both. That's more common than either camp admits.

What each one is

DuckDB is an in-process analytical database — the usual shorthand is "SQLite for analytics". You import it into Python, Node, Go or the CLI, and it runs a columnar query engine inside that process. It reads Parquet, CSV and JSON directly, including from S3, with no load step. There is no server, no port, no cluster. Only one process may open a database file read-write, and while it does, no other process can open the file at all; any number of processes may read it if none is writing. The pattern that scales is one writer producing Parquet and many read-only readers over it.

ClickHouse is a client-server columnar database built for real-time analytics. You run it as a service (or buy ClickHouse Cloud), clients connect over the network, and it stores data in its MergeTree engine: rows land in sorted parts on disk, and background merges combine them. It is built to ingest continuously and answer aggregations over billions of rows in milliseconds, for many clients at once. It scales out across machines when one isn't enough.


This is the first part. The full post — including the rest of the working details — is on my site: DuckDB vs ClickHouse: which one, and when you need both

Top comments (0)