Every Iceberg table accumulates structural debt from its first commit. Files fragment. Snapshots pile up. Manifests bloat. Orphan files sit on object storage that no snapshot references and no query will ever read, but your bill counts every byte. Query performance degrades gradually, then suddenly — and because the architecture is deliberately decoupled, no single engine or catalog owns the operational picture. Nobody is responsible for the health of the lake as a whole.
The format is not the problem. Iceberg ships the primitives: rewrite_data_files, expire_snapshots, remove_orphan_files, rewrite_manifests, compute_table_stats. These are fsck and defrag, not an operating system. They require external intelligence to decide which tables need attention right now, what parameters to use, in what order to run, how to avoid conflicts with the streaming writer committing every 30 seconds, and how to adapt when access patterns shift from one quarter to the next. The missing piece is not a tool. It is a layer.
At 20 tables, you solve that with a script. At 200, the scripts are a full-time job. At 2,000, the scripts need their own team — and that team is no longer building data pipelines.
A lakehouse control plane such as LakeOps exists to close this gap. It sits above the stack you already have — your catalogs (Glue, Polaris, Nessie, Gravitino, Lakekeeper, S3 Tables), your Iceberg tables on object storage, and your query engines — and supplies the operational loop the open components do not provide as one system: sense every table's health through continuous metadata analysis and cross-engine telemetry, decide what maintenance and physical layout each one needs, execute it in the correct dependency order on a dedicated engine, and feed the result back into the next decision.
This article covers why that layer is necessary, what it actually contains, what breaks without it, and how the current landscape of tools, platforms, and managed services fills or fails to fill the gap. For deep dives on individual concerns, see the companion articles on table maintenance, observability, and compaction strategies.
What Iceberg Ships and What It Does Not
It is worth being precise about what Iceberg's own maintenance surface covers, because the gap is both narrower and wider than most people assume.
The maintenance primitives (Iceberg 1.11.0)
| Procedure | What it does |
|---|---|
rewrite_data_files |
Merges small files, applies sort or Z-order, resolves position deletes. 15+ tuning parameters including target-file-size-bytes, min-input-files, partial-progress.enabled, and delete-file-threshold. Branch support added in 1.11.0. |
rewrite_manifests |
Consolidates manifest files for faster query planning. 1.11.0 added sort_by to cluster manifest entries by partition transform rather than accepting default spec order. |
rewrite_position_delete_files |
Compacts position delete files and removes dangling deletes. |
expire_snapshots |
Removes old snapshots and their exclusively-referenced data files. |
remove_orphan_files |
Deletes files on storage not referenced by any live snapshot. Safety threshold: 3-day default minimum age. |
compute_table_stats |
Generates Puffin NDV sketches for cost-based optimization. |
compute_partition_stats |
Calculates partition-level statistics. Incremental from the last stats snapshot. |
Trino's Iceberg connector adds optimize, optimize_manifests, expire_snapshots, and remove_orphan_files as ALTER TABLE ... EXECUTE commands, with hard minimum-retention floors that prevent accidental deletion. Flink embeds maintenance directly in streaming jobs through IcebergSink, with configurable triggers on commit count, file count, file size, and delete-file accumulation.
What none of them provide
Every one of these is a tool, not a system. The gap is not in execution capability. It is in the intelligence that decides when and how to use them:
- No scheduling. No built-in interval, event, or condition-based trigger.
- No health detection. No assessment of whether a table needs maintenance, how urgently, or what kind.
- No cross-engine telemetry. No awareness of query patterns from Trino, Spark, Athena, DuckDB, Snowflake, or Flink. Sort-order selection is guesswork.
- No sequencing. No enforced dependency order. Compacting before expiring snapshots wastes compute on files about to be garbage-collected. Rewriting manifests before compaction produces metadata that goes stale immediately.
- No conflict coordination. OCC retries exist, but nothing coordinates maintenance with concurrent writers to avoid the conflict in the first place.
- No fleet management. No multi-table prioritization, no resource allocation across a lake of hundreds or thousands of tables.
- No policy model. No hierarchical defaults, no inheritance, no configuration that scales beyond one table at a time.
Apache Polaris — now an Apache Top-Level Project — illustrates the gap precisely. Since version 1.0, Polaris defines maintenance policy types: system.data-compaction, system.metadata-compaction, system.orphan-file-removal, system.snapshot-expiry, attachable at catalog, namespace, or table level with inheritance. But — and this is the critical detail most articles miss — Polaris does not execute them. The catalog stores the policy. Who runs it, when, and how is explicitly left to the user. This is a deliberate architectural choice and the right one for a catalog, but the gap between policy declaration and policy enforcement is exactly where a table management platform lives.
What Actually Breaks at Scale
These are not hypothetical concerns. Each is a pattern that emerges between 200 and 2,000 tables — not from any single tool failing, but from the gaps between tools.
Small files make queries exponentially worse
A table with the same data stored in 500,000 × 5 MB files performs 10–50× worse on scan planning, costs 10–100× more in S3 GET requests, and takes dramatically longer to maintain than 5,000 × 500 MB files. In lab benchmarks, compacting 2,500 small files into properly sized ones improved query time from 3.6 seconds to 0.4 seconds for aggregations (8.4× improvement) and from 1.2 seconds to 96 milliseconds for point lookups (12.7× improvement). A 1 TB table costs roughly $0.25 to scan when well laid out and $5.00 when fragmented — a 20× difference in S3 API spend alone, before you count the compute time.
Diagnose it:
SELECT partition,
COUNT(*) AS file_count,
AVG(file_size_in_bytes) / (1024*1024) AS avg_mb,
SUM(CASE WHEN file_size_in_bytes < 8388608 THEN 1 ELSE 0 END) AS small_files,
SUM(file_size_in_bytes) / (1024*1024*1024) AS total_gb
FROM prod.db.events.files
GROUP BY partition
HAVING small_files > 100
ORDER BY small_files DESC;
If a significant fraction of partitions show hundreds of files under 8 MB, every query against that table is paying a planning and API tax that compaction eliminates immediately. Streaming tables with per-minute commits are the most common source, but even daily batch tables fragment when they receive many small appends throughout the day.
Manifest bloat inflates planning time
Planning cost scales with the number of manifests, not the number of partitions. A table with 50,000 partitions and 20 manifests plans quickly. A table with 50 partitions and 20,000 manifests — which is what a per-minute streaming writer produces in two weeks — plans slowly, because the planner must open and parse every manifest in the snapshot before execution begins. Five hundred manifests tracking 200,000 data files requires hundreds of object-store GET requests before any data is read.
Iceberg 1.11.0 added server-side scan planning in the REST catalog spec, which moves manifest parsing from the engine to the catalog, and the sort_by parameter for rewrite_manifests, which clusters manifest entries by partition transform. Both help. Neither removes the need to keep the manifest tree healthy — they make a healthy tree faster and a bloated tree less catastrophic, but the underlying growth still requires periodic rewriting.
SELECT COUNT(*) AS manifest_count,
SUM(added_data_files_count + existing_data_files_count) AS total_entries
FROM prod.db.events.manifests;
When manifest_count exceeds 500 and is growing, planning latency is measurably degraded.
Orphan files silently consume 25–40% of storage
Industry analyses consistently find that 25–40% of object-storage spend on typical production lakes covers bytes no query will ever read again: orphaned files from failed writes, superseded files from compaction, metadata from expired snapshots that nobody cleaned up. One lab measurement on a modest table found 127 MB of orphan metadata against 48 MB of actual data — a 2.8× metadata-to-data ratio. At Databricks, Predictive Optimization has vacuumed exabytes of unreferenced data, saving tens of millions of dollars — a good indication of how much accumulates when expiry does not keep pace with writes.
Delete files impose read amplification
Position delete files in Iceberg v2 require the engine to join each data file with its deletes at read time. Published benchmarks show that once even 3% of rows go through delete-and-insert cycles, query performance drops by roughly 1.5× purely from the file explosion, and the cost is dominated by the fixed overhead of having any delete files at all, not by the number of deleted rows.
Iceberg v3 introduced deletion vectors (DVs), which replace position delete files with a per-data-file bitmap. DVs reduce the read-time join cost dramatically but do not eliminate the storage and statistics impact of keeping deleted rows in data files. The rows stay on storage, consuming space and widening column min/max bounds (which degrades pruning), until compaction rewrites the affected files. DVs make the read penalty cheaper; compaction is still the mechanism that removes deleted data.
SELECT content,
COUNT(*) AS file_count,
SUM(file_size_in_bytes) / (1024*1024) AS total_mb
FROM prod.db.events.files
WHERE content IN (1, 2)
GROUP BY content;
Content 1 = position deletes, content 2 = equality deletes. Any non-trivial count against high-frequency update tables means compaction with delete resolution is overdue.
Snapshots compound everything
A streaming table committing every minute accumulates over 4,000 snapshots per month. Each snapshot extends the metadata chain that the planner walks. Unreferenced data files pile up behind expired snapshots that nobody removed. A table with 90 days of snapshots and no expiry can have 10× more metadata files than data files — and metadata reads are per-query overhead that no amount of caching eliminates for cold starts.
Maintenance storms degrade the lake they exist to fix
IOMETE's engineering team documented a subtler failure mode: at hundreds of tables, the maintenance workload itself degrades the lake. A cron job that compacts 500 tables at midnight generates a burst of metadata commits. Conflict rates spike, retries multiply, and the catalog becomes a bottleneck for production readers and writers that have nothing to do with maintenance.
The irony is structural: maintenance exists to improve query performance, but uncoordinated maintenance at scale degrades it. Production writers hit CommitFailedException because a compaction job holds a stale base snapshot. The compaction job retries, conflicts again, and eventually succeeds — but the writer's commit also retried, potentially doubling data. Or the compaction job succeeds, but the writer committed to a partition that compaction just rewrote, and now the compacted files are immediately fragmented again. Every CommitFailedException represents wasted compute, and at scale the waste is dominated by maintenance-vs-maintenance and maintenance-vs-writer conflicts, not by the actual compaction work.
Physical Layout: Why Binpack Is Not Enough
This distinction is where most Iceberg maintenance articles lose the plot, and it matters enough to state plainly: compaction fixes file count. It does not fix data skipping. If you bin-pack 5,000 small files into 50 correctly sized ones, you reduced planning overhead and API costs. But if every resulting file contains values from the full range of every column, no engine can skip any of them based on a WHERE clause. You have fewer files, but you still scan all of them.
The Iceberg scan planner discards files when column min/max bounds prove a file cannot contain matching rows. That only works when rows with similar values live in the same files. Getting them there requires clustering — and clustering is a separate operation from compaction, with different mechanics, costs, and decisions.
The pruning stack
| Layer | Mechanism | What it decides |
|---|---|---|
| Partition pruning | Partition spec + hidden transforms | Which groups of files can be ignored entirely |
| File pruning (clustering / sort order) | Manifest-level and file-level column min/max bounds | Which files within a partition can be skipped |
| Z-order / multi-dimensional clustering | Space-filling curve bit interleaving | Balanced pruning across 2–4 equally important filter columns |
| Row-group pruning | Parquet footer column statistics | Which row groups within a file can be skipped |
| File size | Target 128–512 MB for analytical tables | Planning overhead vs pruning granularity |
Partitioning decides which groups of files exist. Clustering decides whether the files inside those groups are skippable. These are different operations at different layers. Changing a partition spec is a metadata-only operation. Changing a sort order requires rewriting every data file in scope.
Sort order vs Z-order
A linear sort on customer_id produces files where each file covers a narrow range of customer IDs. Queries filtering on customer_id skip most files. Queries filtering on event_type skip nothing, because event_type values are scattered uniformly across the sorted files.
Z-order interleaves bits from multiple columns so that rows close in multi-dimensional space end up in the same files. This gives balanced pruning across 2–4 filter columns, at the cost of slightly worse pruning on any single column compared to a linear sort on that column. Past 4 columns, Z-order dilutes rapidly, and past 6 it provides little benefit over random distribution.
The decision between them depends on query patterns:
| Workload | Layout |
|---|---|
| Queries overwhelmingly filter on one column | Linear sort on that column |
| Queries filter on 2–4 columns with roughly equal frequency | Z-order on those columns |
| Queries filter on many different columns unpredictably | Z-order on the 3–4 most common; accept limited pruning on the rest |
| High-cardinality point lookups on one key | Linear sort + Parquet bloom filter on that key |
Why this requires telemetry
The right layout depends on how queries actually access the table — and that changes. A table sorted by order_date serves dashboard queries well. When a new fraud-detection pipeline starts filtering on merchant_id, the sort order becomes irrelevant for that workload. Without cross-engine telemetry aggregating filter, join, and group-by frequency across Trino, Spark, Athena, Snowflake, and Flink, sort-order decisions are intuition that nobody revisits.
This is the specific argument for a platform-level solution rather than per-table tooling. The optimal layout is not a property of the table. It is a property of the relationship between the table and every engine that reads it.
What "Table Management" Actually Means
The term gets used loosely. For Iceberg tables in production, it decomposes into five operational concerns that must work as a unified system rather than as independent tools. Each concern feeds the others; assembling them from disconnected parts is technically possible and operationally fragile.
1. Observability
You cannot maintain what you cannot measure. Iceberg has no native health dashboards, no cross-engine telemetry aggregation, no alerting. The metadata is there — snapshot history, manifest lists, file-level statistics — but extracting operational intelligence from it means running Spark or Trino SQL against table.snapshots, table.files, and table.manifests, one table at a time.
A management platform turns raw Iceberg metadata into continuous health classification. Every table is scored — Healthy, Warning, or Critical — based on signals that directly affect query performance: file count and size distribution, manifest depth, snapshot accumulation, delete-file ratio, partition skew, and sort-order alignment with actual query patterns.
The word continuous is the key. A quarterly audit catches problems months after they started costing money. Continuous scoring surfaces a streaming table's small-file explosion within hours, before downstream dashboards slow to a crawl.
Cross-engine telemetry is the other half. When Trino, Spark, and Snowflake all read the same table, understanding which columns each engine filters, joins, and groups on — aggregated into a single access profile per table — is the foundation for layout optimization. Without it, sort-key selection is intuition that nobody revisits.
2. Autonomous maintenance
The maintenance lifecycle for an Iceberg table is a dependency chain, not a checklist:
- Expire snapshots — removes metadata references to old snapshots. This must go first because it shrinks the file set every subsequent step has to process.
- Remove orphan files — deletes data files on storage that no snapshot references, including those just freed by expiry.
- Compact data files — merges small files into optimally sized ones, applying sort or Z-order clustering.
- Rewrite manifests — consolidates manifest files so the metadata tree indexes the layout that now exists.
- Refresh statistics — recomputes Puffin NDV sketches so optimizer estimates match the current file set.
The ordering matters. Compacting before expiring wastes compute — you are rewriting files that expiry would have garbage-collected. Rewriting manifests before compaction produces metadata that goes stale as soon as compaction changes the underlying files. Refreshing statistics before compaction produces estimates that describe a layout that no longer exists.
Most teams start with Airflow DAGs or cron jobs calling Spark procedures. The architectural limitations emerge as the fleet grows:
- Fixed schedules cannot adapt to table state. A table committing every 30 seconds needs compaction multiple times an hour. A weekly batch table needs it once. A uniform daily schedule over-compacts quiet tables and under-compacts hot ones.
- Cross-DAG coordination is absent. Ensuring expiry completes before compaction starts, across 500 tables, requires manual orchestration Airflow was not designed for.
- Commit conflicts accumulate. Compaction jobs running alongside production writers compete for the metadata pointer via OCC. At scale, conflict rates become the dominant source of failed maintenance, wasted compute, and on-call pages.
- Linear configuration burden. Each new table needs its own sort columns, target file size, retention window, and schedule. At 2,000 tables, maintaining the configurations is a full-time job.
3. Intelligent compaction and layout
Standard bin-pack compaction merges small files into bigger ones. As established above, it does not improve data skipping. Query-aware compaction goes further: it builds a per-table access profile from all connected engines' query logs, identifies the filter, join, and grouping columns that drive data access, and physically re-sorts data files to match during compaction. When data is sorted by the columns queries actually hit, Parquet row-group min/max statistics become tight, and engines skip entire file groups without reading them. The difference is often an order of magnitude: a table that scanned 4 TB scans 400 GB after sort compaction, because min/max pruning eliminates 90% of files.
Layout simulation is the related capability: testing a proposed sort strategy against real query patterns on an Iceberg branch before committing it to production. Bad sort decisions are expensive to reverse since you are rewriting all the data files in scope, so the ability to evaluate projected scan reduction before a full rewrite is not a nice-to-have. It is the mechanism that makes automated sort-order selection safe rather than reckless.
4. Multi-engine coordination
Production Iceberg lakehouses rarely run a single engine. Trino handles interactive analytics. Spark runs ETL. DuckDB serves quick lookups. Snowflake powers BI. Athena covers ad-hoc. Flink processes streaming ingestion. Each has different cost characteristics, latency profiles, and strengths. Without coordination, every query goes to one engine regardless of whether it is the right one, and maintenance decisions cannot account for how all engines actually access the data.
A management platform provides a routing layer that dispatches each query to the optimal engine based on query shape, cost model, and latency target. This creates a reinforcing loop with compaction: a well-sorted table enables cheaper engines to serve queries that a fragmented table would require an expensive engine for. Routing telemetry feeds back into sort-order selection, so both improve simultaneously.
5. Policy and governance
At 50 tables, you configure each one individually. At 500, you need declarative policies that cascade hierarchically — organization, catalog, namespace, table — with inheritance and overrides. Define compaction targets, snapshot retention, orphan cleanup thresholds, and sort strategies once at the namespace level. New tables inherit the correct configuration automatically. Override individual tables when needed.
Governance also means audit trails. Every maintenance operation — what ran, when, duration, files before and after, bytes reclaimed, the signal that triggered it — needs to be logged in a structured, queryable format. For compliance environments the trail is evidence; without it, you are proving GDPR erasure completion from Airflow logs and Spark driver output.
Why these five must be unified
Each concern has standalone tools. The problem is that they are deeply coupled in practice:
- Observability feeds maintenance. Health scores determine which tables need attention and how urgently.
- Maintenance feeds compaction. Expiry and cleanup must complete before compaction to avoid rewriting files about to be garbage-collected.
- Compaction depends on telemetry. The optimal sort order comes from knowing how all engines query the table.
- Routing improves when compaction improves. A well-sorted table enables cheaper engines, and routing telemetry feeds the next sort decision.
- Policies must span all four. A retention policy affects snapshots (maintenance), which affects orphan accumulation (observability), which affects how much compaction processes, which affects file layout for every engine (routing).
The argument for a platform is not that individual tools are bad. It is that the interactions between concerns are where the operational value compounds, and independent tools by definition cannot exploit feedback loops across them.
A Diagnostic Runbook: Is Your Lake Degraded?
Before choosing a tool, measure the problem. Run these queries against your most important tables and score the results.
Step 1: File fragmentation
SELECT COUNT(*) AS total_files,
COUNT(CASE WHEN file_size_in_bytes < 8388608 THEN 1 END) AS small_files,
AVG(file_size_in_bytes) / (1024*1024) AS avg_mb,
STDDEV(file_size_in_bytes) / (1024*1024) AS stddev_mb
FROM prod.db.events.files;
Healthy: small_files < 5% of total_files, avg_mb in the 128–512 MB range. Warning: 5–30% small files. Critical: > 30% small files, or avg_mb under 32.
Step 2: Manifest depth
SELECT COUNT(*) AS manifest_count,
AVG(length) / 1048576 AS avg_manifest_mb,
SUM(added_data_files_count + existing_data_files_count) AS total_entries
FROM prod.db.events.manifests;
Healthy: under 200 manifests. Warning: 200–500. Critical: over 500 and growing.
Step 3: Snapshot accumulation
SELECT COUNT(*) AS snapshot_count,
MIN(committed_at) AS oldest_snapshot,
MAX(committed_at) AS newest_snapshot
FROM prod.db.events.snapshots;
Healthy: retention window is enforced and recent. Warning: snapshots older than 30 days with no policy. Critical: thousands of snapshots spanning months.
Step 4: Delete file burden
SELECT content, COUNT(*) AS files, SUM(record_count) AS rows
FROM prod.db.events.files
WHERE content > 0
GROUP BY content;
Any equality deletes outside a Flink pipeline are a red flag. Position delete counts exceeding 10% of data file count indicate compaction lag.
Step 5: Sort-order alignment
SELECT readable_metrics.customer_id.lower_bound AS lb,
readable_metrics.customer_id.upper_bound AS ub,
file_size_in_bytes / (1024*1024) AS mb
FROM prod.db.events.files
ORDER BY lb
LIMIT 20;
Replace customer_id with your primary filter column. If adjacent files have broadly overlapping bounds (lower_bound = 1, upper_bound = 10M on most files), the table is not clustered on that column and no amount of partitioning will fix file-level pruning.
Scoring
| Score | Meaning | Action |
|---|---|---|
| All green | Table is well-maintained | Monitor; adjust cadence if write patterns change |
| 1–2 warnings | Targeted fix | Run the relevant procedure(s); consider policy-based scheduling |
| Any critical | Active degradation | Immediate maintenance; evaluate platform-level management |
| Multiple criticals across tables | Systemic gap | The maintenance approach itself needs replacing |
The Current Landscape
The Iceberg management space has converged on several architectural approaches, each making different trade-offs about scope, lock-in, and operational burden.
Catalog-bundled maintenance
| Platform | What it provides | What it does not |
|---|---|---|
| Apache Polaris (TLP since Feb 2026) | Maintenance policies (data-compaction, metadata-compaction, orphan-removal, snapshot-expiry) with catalog-namespace-table inheritance | Does not execute them. Policy declaration without enforcement. |
| Dremio Open Catalog (on Polaris) | Automatic compaction, DV resolution, clustering, manifest rewrite, vacuum — five operations, ~3h optimize / 24h vacuum, dedicated engine | Only tables inside Dremio's catalog. Target file size not per-table configurable. |
| AWS Glue | Three optimizers: compaction (binpack/sort/z-order), snapshot retention, orphan deletion. $0.44/DPU-hour. | No manifest rewrite, no statistics, no cross-engine sort awareness, no sequencing, Parquet only, auto-suspends after 4 consecutive failures. |
| S3 Tables | Fully managed compaction (auto/binpack/sort/z-order), snapshot management, orphan cleanup — all enabled by default | No manifest rewrite, no statistics, no cross-engine telemetry. Object monitoring fee penalizes the fragmentation compaction exists to fix. |
| Snowflake | Data compaction (bin-pack) + manifest compaction (always-on, cannot disable). Free for Snowflake-only writes since May 2026. | No sort/z-order beyond Automatic Clustering. External-engine DML activates compaction credit charges. No exposed expire_snapshots or remove_orphan_files procedures. |
| Databricks | Predictive Optimization: OPTIMIZE + Liquid Clustering + VACUUM + ANALYZE, enabled by default since Nov 2024. Selects and evolves clustering keys from query telemetry. | Requires Unity Catalog. Telemetry is scoped to Databricks workloads. |
| Starburst Galaxy | "Icehouse LakeOps": automated compaction (hourly), snapshot expiry (daily, >30 days), manifest rewrite (daily), orphan removal. Per-table health dashboards. | Schedule-driven, not health-driven. No cross-engine sort awareness. |
| Lakekeeper | Automatic snapshot expiry (per-warehouse, configurable) and adaptive orphan cleanup | No compaction, no manifest rewrite, no layout optimization. |
| Gravitino | Table Maintenance Service: alpha, policy-driven compaction for identity-partitioned tables only | CLI-driven, not self-scheduling. No snapshot expiry, orphan cleanup, or sort maintenance. |
| Nessie | None | Pure catalog: Git-style branching and tagging for metadata. |
The pattern: each platform solves maintenance within its boundary and is blind to everything outside it. Databricks is the furthest along with query-driven clustering, but it requires Unity Catalog and sees only Databricks workloads. In a multi-engine, multi-catalog environment — which is what the open lakehouse was designed to enable — each boundary creates a gap.
Open-source management tools
| Tool | Language | Stars | What it does | Limitation |
|---|---|---|---|---|
| Apache Amoro (incubating) | Java | ~1,170 | Self-optimizing management service (AMS) with optimizer groups. Monitors tables, dispatches compaction, supports Iceberg + Paimon + Hudi. | AMS calls loadTable() on every planning cycle (default: every minute). At 2,000+ tables, this creates significant catalog and object-storage pressure. Still incubating (v0.8.1). |
| OpenHouse | Java | ~400 | LinkedIn's RESTful control plane: declarative table provisioning, policy-based retention, Spark maintenance on Kubernetes. Proven at LinkedIn scale. | Built on Iceberg 1.2.0 and Spark 3.1.2 — multiple major versions behind current Iceberg. Tied to LinkedIn's infrastructure patterns. |
| Nimtable | TypeScript | ~479 | Web platform for Iceberg observability and management by RisingWave Labs. Browse catalogs, visualize file and snapshot distribution, run SQL, acts as REST Catalog endpoint. | Observability-focused; delegates maintenance execution to Spark or RisingWave. Not a compaction engine. |
| Floe | Java | ~29 | Policy-based maintenance orchestrator. Declarative policies with glob-pattern table matching, health thresholds, scheduled execution of all four Iceberg procedures via Spark (Livy) or Trino. Multi-catalog. Web UI. Featured on Apache Polaris blog as the reference policy consumer. | Orchestrator, not engine — dispatches to Spark/Trino for execution. Solo developer. |
| nimtable/iceberg-compaction | Rust | ~140 | Purpose-built Rust compaction on DataFusion. Handles positional and equality deletes. Most active open-source compaction project. | Full-table bin-pack only; sort/Z-order, incremental compaction, and orphan deletion are roadmap items. |
| Bergman | Rust | 1 | Rust maintenance engine on iceberg-rust + DataFusion. All four operations plus dangling delete-file removal. Library-first. |
Pre-release, solo project. Implements its own commit layer because upstream iceberg-rust lacks rewrite-commit support. |
| Firn | Go | ~15 | JVM-free daemon; shells out to DuckDB for compaction. Binpack/sort/z-order, snapshot expiry, orphan cleanup. Multi-catalog (Lakekeeper, Glue, Polaris, Nessie). | Pre-v1.0 and stale for several months. |
| Zamboni | Python | — | Single-process maintenance via PyIceberg + DuckDB. Declarative layout config. Covers compaction, Z-order, partition evolution, manifest rewriting, snapshot expiry, orphan removal. | Single-process; no fleet orchestration, no health-triggered scheduling. |
| IceGuard | TypeScript | 3 | Multi-catalog web console for browsing and scheduling maintenance pipelines (Airflow-style chained actions). | Very early stage. |
A few deserve deeper examination because they represent distinct architectural philosophies:
Apache Amoro is the most ambitious open-source effort. Its architecture (AMS → Optimizer Groups → Optimizer containers) provides fleet-level visibility: a central dashboard shows every table's health, a scheduling layer assigns compaction tasks to optimizer groups, and the system detects "self-optimizing" conditions (small files, orphan snapshots, excessive delete files). The architectural concern is scale: at 2,000+ tables against S3-backed catalogs, the per-minute metadata load from loadTable() calls becomes the bottleneck the system was designed to alleviate.
OpenHouse proves the concept at LinkedIn scale. It is a full RESTful control plane: declarative table provisioning, policy-driven retention, column-lineage-based governance, Spark maintenance jobs on Kubernetes, and an API-first design where tables are requested through a service and automatically provisioned with the right policies. The limitation is that it was built for LinkedIn's specific infrastructure and remains on Iceberg 1.2.0 — five major versions behind.
nimtable/iceberg-compaction has emerged as the most active open-source compaction engine. Written in Rust on DataFusion, it handles positional and equality deletes, runs without a JVM, and benchmarks competitively. Its scope is intentionally narrow: full-table bin-pack compaction. This is a compaction library, not a management platform, and it is useful precisely because of that focus.
What none of them provides is the full loop: health detection driving prioritization, dependency-aware sequencing, cross-engine telemetry feeding layout decisions, conflict-safe execution, and a feedback cycle where outcomes improve future decisions. Each solves a slice.
Where Flink fits
Flink's TableMaintenance is the most sophisticated per-table automation in Iceberg: event-driven triggers on commit count, file count, file size, and delete-file accumulation, all embedded in the streaming topology. But it is per-table, per-sink maintenance — not a management platform. No fleet-wide visibility, no cross-engine telemetry, no policy management, no multi-table prioritization. It is one component of what a platform provides.
How a Control Plane Runs This Loop
A table management platform occupies a specific architectural layer:
┌──────────────────────────────────────────────────┐
│ Query Engines (Trino, Spark, Flink, DuckDB...) │
├──────────────────────────────────────────────────┤
│ TABLE MANAGEMENT PLATFORM / CONTROL PLANE │ ← This layer
│ Health · Sequenced maintenance · Telemetry · │
│ Layout optimization · Policies · Routing │
├──────────────────────────────────────────────────┤
│ Catalog (Polaris, Nessie, Glue, Gravitino...) │
├──────────────────────────────────────────────────┤
│ Table Format (Apache Iceberg) │
├──────────────────────────────────────────────────┤
│ Object Storage (S3, GCS, ADLS) │
└──────────────────────────────────────────────────┘
It does not replace the catalog, the engine, or the storage. It sits between them and supplies the operational intelligence that none of them individually possesses. A catalog manages metadata and access control. A query engine executes queries and maintenance procedures. A control plane decides which maintenance to run, when, in what order, with what parameters, and adapts those decisions continuously.
LakeOps is built for this layer. It connects to catalogs through standard APIs — Glue, Polaris, Nessie, Gravitino, Lakekeeper, S3 Tables — and to query engines for telemetry, without requiring data movement, pipeline changes, or an SDK. The architecture follows a closed-loop model:
Sense — reads Iceberg metadata from every connected catalog continuously. Each table is scored as Healthy, Warning, or Critical based on file layout, manifest depth, snapshot age, delete-file ratio, partition skew, and sort-order alignment with observed query patterns. Cross-engine telemetry from Trino, Spark, Snowflake, Athena, DuckDB, and Flink reveals which columns are filtered, joined, and grouped — per table, across the full engine fleet.
Plan — prioritizes tables by degradation severity. A production table with 50,000 small files gets compacted before a staging table with 200. Operations are sequenced in dependency order. Partitions with active writers are identified and excluded, eliminating the commit conflict failures that dominate maintenance at scale.
Optimize — compaction, snapshot expiration, orphan cleanup, manifest rewriting, and delete-file resolution execute on a Rust engine built on Apache DataFusion. No Spark clusters, no JVM overhead. Maintenance operations are staggered across the catalog to avoid the maintenance-storm pattern described earlier.
Learn — outcomes feed back into the system. Tables where sort compaction produced large scan reduction get prioritized for sort maintenance. Sort orders adapt as query patterns evolve. Cadence tunes to each table's actual write velocity. This is the closed loop that independent tools cannot form: the one where every maintenance decision improves the next one.
What this looks like concretely
Sort-order selection from telemetry. Rather than a human analyzing logs across three engines to guess which columns to sort by, the platform aggregates WHERE, JOIN, and GROUP BY column frequency across every connected engine and builds a per-table field-access profile. If 82% of Trino queries filter on user_id and 71% of Athena reports filter on campaign_id, the sort order maximizes pruning for that mix. When the mix changes — a new dashboard, a new pipeline — the sort order adapts in the next compaction cycle.
Layout simulation on Iceberg branches. Before a sort-order change touches production data, the candidate layout is tested against real query patterns on an Iceberg branch: comparing projected scan reduction and file distribution to the baseline. Only a strategy that measurably improves pruning gets applied. Bad sort decisions are expensive to reverse (you are rewriting every data file in scope), and this mechanism is what makes automated layout selection safe rather than reckless. Compaction strategy details cover how binpack, sort, and Z-order selection differ per table.
Conflict-aware compaction. Rather than failing four times and suspending (as Glue's optimizer does), the platform monitors per-partition commit activity across engines and excludes partitions with active writers from the compaction scope. When a commit does lose an OCC race, it backs off, re-evaluates partition state, and reschedules rather than immediately retrying into the same conflict.
Hierarchical policies. Maintenance configuration cascades from catalog to namespace to table, with more specific policies overriding broader defaults. New tables inherit the right configuration from their namespace automatically. This is the same model Polaris's policy framework uses for declaring intent, with the crucial difference that a control plane actually executes it.
Staged adoption. You can start in manual mode, executing operations per table and reviewing results. Then define scheduled policies at the namespace level. Then enable Adaptive Maintenance, where the platform monitors health signals and runs the right operations at the right time without fixed schedules or manual tuning.
Performance. The Rust/DataFusion compaction engine avoids the JVM entirely. Published benchmarks show 95% faster compaction than Spark on identical datasets: a 200 GB S3 Tables compaction completed in 221 seconds vs. 1,612 seconds for Spark (7.3×), while Spark OOMed on a 1.2 TB table that the Rust engine processed in approximately 11 minutes. A 5.5 TB, 10-table benchmark achieved 2,522 MB/s peak throughput with 81% file-count reduction. Sub-minute compaction cycles keep up with streaming writers instead of running as batch catch-up jobs.
The platform overview covers how these capabilities compose: from lake-wide observability through query-aware compaction, multi-engine query routing, declarative policy management, and AI-agent access with layered guardrails (ReadOnly mode, cost-estimate caps, PII masking, human approval for high-stakes operations).
Three architectural guarantees matter for teams evaluating this category:
- No data movement. All operations happen through standard Iceberg commit APIs. Data stays on your storage.
- No pipeline changes. Connect catalogs and engines through standard APIs. No SDK, no code changes.
- No lock-in. Everything the platform produces is standard Iceberg: snapshots, manifests, data files. Disconnect it and your tables remain exactly as they are.
Evaluating the Options
The right approach depends on where you are, not where you want to be:
| Your situation | Approach |
|---|---|
| Under 50 tables, single engine, single catalog, batch writes | Scripts + standalone tools (Spark procedures, Trino EXECUTE) |
| Glue catalog, Athena-only, Parquet, batch, and compaction works fine | AWS Glue optimizers. Add catalog-level defaults and stop there. |
| Glue optimizer suspends on streaming writes or manifests are never rewritten | Supplement Glue with external tools over the REST catalog for the missing operations |
| Single platform (Databricks, Snowflake) genuinely handles 90%+ of your workload | Platform-bundled optimizer. Accept the boundary. |
| Multiple engines, multiple catalogs, or streaming writes with cross-engine access | Control plane. The integration cost of assembling standalone tools exceeds the cost of a platform. |
| 500+ tables, growing, with maintenance consuming more engineering time than pipelines | Control plane. Manual approaches break at this scale. |
Build it yourself (Airflow + Spark)
Works when you have under ~100 tables, a single engine, and a platform team with capacity to own the scheduler, conflict handling, sequencing, health detection, and configuration management. The honest failure mode is not that it does not work. It is that it works until the person who built it changes teams.
One detail worth knowing: Iceberg's min-input-files default is 5, while Glue overrides it to 100. If you are moving from Glue compaction to Spark procedures, that difference explains why your tables feel under-compacted on Glue.
Adopt a catalog-bundled optimizer
Works well when a single engine or platform is genuinely your primary query layer. Databricks Predictive Optimization is the most complete, with query-driven clustering key selection that evolves over time, but it requires Unity Catalog and sees only Databricks workloads. Snowflake's managed compaction is zero-config but bills you permanently the moment an external engine touches the table. Glue is the path of least resistance on AWS but cannot rewrite manifests, select sort orders, or survive streaming write conflicts. Each is useful within its scope and blind outside it.
Use open-source management tools
Floe gives you policy-driven scheduling over Spark or Trino with multi-catalog reach. nimtable/iceberg-compaction gives you a fast Rust compaction engine. Apache Amoro gives you a self-optimizing management service with streaming-aware triggers. These are real options for teams that want to own the infrastructure and avoid vendor dependency, with the honest caveat that none provides the full loop and most are pre-1.0.
Adopt a dedicated control plane
Works when the table fleet is large enough that per-table configuration is intractable, when multiple engines hit the same tables and need coordinated layout decisions, when streaming writes make conflict handling a constant concern, and when the engineering time spent maintaining the maintenance system exceeds the cost of a platform. This is where LakeOps fits: a managed control plane that layers onto your existing stack, runs the full operational loop, and scales sublinearly — going from 500 to 5,000 tables should not require 10× more configuration or human attention.
The hybrid path
These approaches compose more than most comparisons suggest. You can run Glue compaction for bulk binpack on the tables it handles well, connect an open-source tool for manifest rewriting, and add a control plane for the tables with streaming writes and multi-engine access patterns. The Glue REST catalog makes this possible; the question is whether the integration cost exceeds the cost of a single system.
How Format Evolution Changes the Picture
Iceberg v3 and the early v4 proposals affect what management platforms need to handle, but they do not reduce the need for management itself.
Deletion vectors (v3) replace position delete files with per-data-file bitmaps. The read-time penalty drops from a file join to a bitmap lookup, but deleted rows stay in data files, consuming storage and widening column statistics until compaction rewrites them. DVs make the read penalty cheaper; compaction remains the mechanism that removes deleted data.
Variant shredding (v3) stores semi-structured data (JSON) in typed Parquet columns while preserving the original document. This creates wider tables with more columns, more partition statistics, and more metadata — making sort-order decisions more complex because shredded field access patterns may differ from top-level columns.
Row lineage (v3) assigns a monotonically increasing sequence number to every row, enabling engines to track individual rows across snapshots. For management platforms, lineage simplifies delete-file resolution: identify exactly which data files are affected by a batch of deletes and compact only those, rather than scanning the entire table.
Geospatial bounds (v3) add bounding-box statistics for geometry columns — a new class of statistics that management platforms need to maintain, and spatial clustering (Z-order on coordinates) interacts with them in ways that temporal sort does not.
Metadata compaction (v4 proposal) would move manifest consolidation into the format itself, potentially eliminating external rewrite_manifests calls. If adopted, this removes one of the five maintenance procedures but does not affect the other four.
The net effect: the individual procedures change — some get simpler, some get more complex, and new ones emerge. The need for an external system to coordinate them does not diminish.
Where the Market Is Heading
The Iceberg management platform category is emerging because the open lakehouse reached a maturity inflection. Iceberg is the format — every major cloud and engine treats it as the default. The catalog layer has consolidated around the REST spec with Polaris, Nessie, and managed offerings. What remains unsolved is the operational layer above catalogs and engines: the layer that keeps tables healthy, queries fast, and costs controlled without requiring every team to become an Iceberg internals expert.
The direction is visible from every side. Databricks acquired Tabular (the company founded by Iceberg's creators to build exactly this managed layer) and absorbed it into Predictive Optimization. Cyera paid $100–130M for Ryft, an Iceberg operations startup with $8M in seed funding — a security vendor paying nine figures for table-level operational intelligence, because governed, traceable data access cannot exist without maintained, well-managed tables underneath. Starburst launched "Icehouse LakeOps" as a named product for automated table health. SAP acquired Dremio, committing to Polaris and Iceberg as the foundation of its Business Data Cloud. Apache Amoro remains in incubation. Open-source projects are multiplying around the gap. And Iceberg v4 proposals include an adaptive metadata tree that would fold manifest consolidation into the format itself, potentially eliminating external rewrite_manifests calls.
The question is not whether Iceberg tables will be managed. It is whether that management intelligence lives inside a proprietary platform, gets assembled from scripts and tools by every team independently, or runs as a dedicated control plane that preserves the openness Iceberg was built to provide.
Getting Started
If you are feeling the operational weight — slow queries traced to small files, storage costs climbing from orphan accumulation, maintenance DAGs consuming more engineering time than the pipelines they were supposed to enable — run the diagnostic runbook above against your most critical tables. Score them against the health criteria. If two or more tables show warnings or critical scores across multiple dimensions, the maintenance approach itself needs attention, not just the individual tables.
LakeOps connects to your existing catalogs and surfaces health classification immediately — often revealing degradation you knew existed but could not quantify.
The adoption path is designed for progressive trust:
Connect a catalog. LakeOps connects to Glue, Polaris, Nessie, Gravitino, Lakekeeper, or S3 Tables through their standard APIs. No code changes, no SDK, no agents to deploy. Immediate visibility into every table's health, file distribution, manifest state, and snapshot history.
Run manual maintenance. Execute individual operations — expire snapshots, compact a single partition, rewrite manifests — and review the results. You control what runs and when.
Define policies at the namespace level. Set compaction targets, retention windows, sort strategies, and maintenance cadence that new tables inherit automatically. Override specific tables when needed.
Enable adaptive maintenance. The platform monitors health signals continuously and runs the right operations at the right time, in the correct dependency order, with conflict-safe execution. Tables that need attention get it immediately; tables that are healthy are left alone.
The lake becomes self-maintaining. The on-call burden shifts from firefighting fragmentation and orphan accumulation to reviewing a structured event log. The platform team gets its capacity back for the work that actually produces value — building data products, not babysitting compaction jobs.


















Top comments (1)