The first time I lost a table to a metadata problem, the data was fine. Every Parquet file was sitting right there in the bucket. The pointer to the current snapshot was wrong, and for around 4 hours the table reported zero rows. Nobody looked at the data; everybody looked at the dashboard.
That's the part of the data lakehouse most people skip when they study. They learn the format names, memorize "Iceberg supports time travel," and move on. In 2026 the argument stopped being about which table format wins. It's now about where the metadata should live: as files on object storage, or as rows in a SQL database. That's what you'll be asked to reason about in your next interview, and you can answer it if you understand how commits, snapshots and catalogs work. Vendor names won't get you there.
DuckLake 1.0 and the case for metadata in a database
DuckLake started as a manifesto. In May 2025, Mark Raasveldt and Hannes Mühleisen (the DuckDB folks) published the original DuckLake post. Its core claim is that metadata is a database problem and doesn't belong in the file system. Less than a year later, on April 13, 2026, DuckLake 1.0 shipped in DuckDB 1.5.2 with a production-ready spec, a reference implementation, and guaranteed backward compatibility.
The design is simple enough to explain on a whiteboard in 60 seconds, so practice doing exactly that. Schemas, file locations, snapshots and the transaction log all live in a catalog database: Postgres if you're self-hosting, SQLite or DuckDB locally, MotherDuck in the cloud. Data stays as Parquet on object storage. To find the files for a query, you run one SQL query with a WHERE clause. MotherDuck's architecture write-up compares that to Iceberg's 3 layers of metadata files, each a sequential read of roughly 100ms.
The headline feature is data inlining. Small inserts and deletes (the default threshold is 10 rows) skip Parquet entirely and get written straight into the catalog database. A background process flushes them to object storage later. That goes straight at the small files problem, which has hurt every streaming-into-the-lake project I've touched.
DataLakeHouse Hub put a number on that problem: streaming 100 rows per second into an Iceberg table, with 100 active partitions and 60-second commits, produces about 144,000 files a day. So you schedule compaction, compaction starts fighting with your writers, and eventually someone pages you.
DuckLake's data inlining post claims big numbers against Iceberg on a 100-second streaming workload: 105x faster inserts, 923x faster aggregation, 189x faster checkpointing. These are DuckLake's own benchmarks on a workload chosen to show off inlining, so read them as a demonstration of the design and don't treat them as an independent bake-off. MotherDuck also cites roughly 100 transactions per second for DuckLake versus about 1 TPS for file-based formats, and a 2026 research paper on SQL-backed catalogs measured median commit latency of 8.4ms against 180ms for Iceberg. The direction is consistent. The exact multipliers depend on who's running the test.
The 1.0 release also included 68 reliability and correctness PRs, sorted tables for range queries, Iceberg-compatible bucket partitioning, GEOMETRY and VARIANT types, and experimental Puffin-based deletion vectors. That last one matters, because it shows DuckLake hedging toward Iceberg V3 compatibility instead of trying to fight it.
It makes me laugh that I spent years watching the industry move metadata out of the Hive metastore, which was a relational database, and onto object storage because the database was "the bottleneck." Now the most-discussed new format puts metadata back into Postgres. I've been through enough cycles to stop being surprised; the hot new thing is usually the old thing with better engineering and a nicer logo.
The economics hold up, though. A managed Postgres instance costs less per month than one afternoon of a senior engineer babysitting compaction jobs. If your workload is lots of small writes, paying for a database to absorb them is cheap next to the engineer time it takes to clean up 144,000 files a day.
Iceberg 1.11, the V3 spec, and what changes for data engineers
Iceberg didn't sit still. Apache Iceberg 1.11.0 landed in May 2026; sources put the date anywhere from May 15 to May 27 depending on whether they mean the vote, the artifacts or the blog post. The Google Open Source announcement covers the main points:
- Java 17 is the new floor. Java 11 is gone.
- Spark 4.1 and Flink 2.1 are the default build targets. Spark 3.4 is deprecated and scheduled for removal in 1.12; Flink 1.19 support is already removed.
- Remote scan planning lets the REST catalog server do the metadata work and stream back file scan tasks.
- Built-in table encryption uses envelope encryption and supports Google KMS.
The bigger story is the V3 spec, ratified in June 2025 and now production-mature in 1.11.
Deletion vectors come first. In V2, frequent deletes generated piles of positional delete files, so readers had to merge them at query time. V3 replaces them with Roaring bitmaps attached to each data file and stored in Puffin files, Iceberg's binary blob container. Multiple deletion vectors share one physical file. AWS benchmarked this on EMR and measured 55% faster deletes, delete files shrinking from 1,801 bytes to 475 bytes, and full scans 28.5% faster. For CDC and MERGE-heavy tables, that's a real improvement.
Variant is a single column that holds values of arbitrary, evolving shape. It's the native replacement for the "dump the JSON into a string column and parse it at query time" workaround all of us have shipped at least once. Shredding (storing Variant fields in columnar form) currently only works with Parquet.
Row lineage adds _row_id and _last_updated_sequence_number, which give every row an identity and a change marker. That's useful for CDC and audit.
The V2 to V3 upgrade is one-way. The ALTER TABLE itself is metadata-only and rewrites no data, so it feels harmless. Once a table writes V3 metadata, though, any engine that only understands V2 can't read it. I've lived through a version bump that broke a reader nobody remembered existed: a quarterly finance extract running on an old cluster, owned by a person who'd left 2 years earlier. We found it when finance called.
Before you upgrade anything:
- Inventory every reader and writer: every engine, every version, every scheduled job. Trace them through query logs, because the docs will be wrong.
- Get Java 17 everywhere first. Iceberg didn't make this choice lightly; maintainers moved after contributors on Java 21 kept hitting build failures against modules pinned to 11. Your Spark runtime, your Flink jobs and your random ingestion service all need it.
- Upgrade a table for a feature, and only for a feature. If you need deletion vectors, Variant or encryption, upgrade that table. If you don't, V2 tables keep working.
- Expect governance to be the hard part. As one DataLakeHouse Hub piece put it, catalog migrations "fail rarely on the format and almost always on governance." Access control and permissions are where the weekend goes.
The catalog is where vendors are fighting now
Look at the business news this year and the pattern is hard to miss: everyone is putting a database back in the middle of the lake.
SAP is buying Dremio. The deal was announced May 4, 2026, with close expected in Q3. SAP says it's "fully committed to continuing to invest in and prioritize" Iceberg, Polaris and Arrow, and plans to build Business Data Cloud into an Iceberg-native lakehouse with an open catalog on Polaris and the Iceberg REST API. Dremio has been one of Iceberg's main stewards. When a company that size buys a core contributor, ask what happens 3 years from now, after the press release has been forgotten.
The safety net is governance. Apache Polaris graduated to a top-level Apache project on February 15, 2026, after 18 months of incubation, 6 releases and over 2,800 merged PRs. The vote passed with 27 binding +1s and no objections, and contributors include Snowflake, AWS, Google Cloud, Azure, dbt Labs, Stripe and others. With that many companies committing code, one acquirer can't take the project over without the others noticing. If you're on Dremio, the Iceberg REST Catalog API is your exit door. Keep it open.
Databricks is putting Postgres inside the lakehouse. It bought Neon for about $1 billion in May 2025, citing that over 80% of Neon databases were being created by AI agents. In August 2026, Databricks reported Lakebase above $100M revenue run-rate, growing twice as fast as its Lakehouse product. In May, Databricks also made Catalog Commits generally available for Unity Catalog managed Delta tables. That routes commits through the catalog so external engines can't write straight to storage and leave the catalog and the files disagreeing.
Snowflake did its own version. It bought Crunchy Data for $250 million and open-sourced pg_lake in November 2025, which lets Postgres create and query Iceberg tables directly.
Iceberg won broad engine support; Delta and Hudi both added Iceberg compatibility. The fight has moved one layer up, to the catalog, because whoever owns the commit owns the table. File-based metadata and database metadata are 2 answers to the same question: who gets to say what the current version of a table is, and how fast can they say it?
Formats are converging. Catalogs are where the lock-in lives now. Pick your catalog like you're going to be stuck with it, because you probably are.
Explaining table formats in a data engineering interview
Interviewing is a separate skill from the job. I've been on panels where a candidate who'd run Iceberg in production for 2 years fumbled "what does the catalog actually do?" because they'd never had to say it out loud. That's what costs people the level. They have the experience; they just can't put it into words under pressure.
These are the concepts that keep showing up, and how I'd answer each.
"What's the difference between Parquet and Iceberg?" Parquet is a file format; it describes how bytes are laid out in one file. Iceberg is a table format: snapshots, atomic commits, schema evolution, partition evolution, and the metadata that tells an engine which files make up the table right now. If you say "Iceberg is Parquet with metadata," a good interviewer will keep pushing until you fall over.
"What does a catalog do?" It answers one question: where is the current metadata for this table? A commit is an atomic swap of that pointer. Two writers race and one wins; the loser retries against the new snapshot. Every catalog design, including Polaris, Unity, DuckLake's Postgres and Glue, is a different way of making that swap fast and safe.
"What isolation do you get?" Snapshot isolation. Readers never block writers, and writers never corrupt an in-progress read. Candidates regularly call this serializable. It isn't, and saying so correctly tells a senior interviewer you've thought about concurrent writers. Bonus points for knowing Iceberg tracks snapshots through manifest lists, Delta through a sequential JSON commit log, and Hudi through a timeline.
"Why would anyone put metadata in a database?" Small, frequent writes. Under CDC or sub-second commits, the bottleneck is metadata overhead, and data I/O stops mattering. Object storage charges around 100ms per request; a database answers in milliseconds and handles concurrency well. That's the reason DuckLake exists, and the reason LSM-based designs like Paimon exist.
"Which format would you pick?" Never answer with a benchmark from a blog post. Name the axes: which engines need to read and write the table, what your team can actually operate, and how much vendor independence you need. For a DuckDB-heavy shop with a modest team, DuckLake on Postgres is operationally light. For 4 engines across 3 teams with real governance requirements, Iceberg behind a REST catalog is the default for a reason.
Then add the opinion that separates seniors from everyone else: most of y'all don't need 100 transactions per second. Batch still covers most of what companies run. The ability to stream into the lake is nice to have; making a 6-hour batch job reliable is what keeps you employed. Say that in the interview, with a reason, and you sound like someone who's carried the pager.
You don't need a vendor certification for any of this. Every concept above (atomic commits, isolation levels, write amplification, metadata lookup cost) existed before these products and will outlive them. A DuckLake catalog is literally tables in Postgres you can query, so SQL fluency pays off twice. If you need to sharpen it, for sql practice problems, DataDriven is great, and we built it as interview prep for the SQL half of the loop. Use it to free up your head for the architecture questions, which is where these conversations get decided.
The tools in this article will look different in 18 months. Somebody will be acquired, somebody will fork something, and Iceberg V4 will promise single-file commits. Commits, snapshots and the catalog pointer will still be what you're debugging at 3am.
So, for those of you already running a lakehouse in production: where does your metadata live today, and has a catalog ever lied to you about what's in a table?
Top comments (0)