Why I chose this topic: Three years ago, I spent an entire weekend debugging a corrupted Hive Metastore (HMS) that locked up our daily reconciliation pipeline. I'm writing this because the industry is finally moving toward decoupled, REST-based cataloging, and I want you to avoid the migration headaches I dealt with by learning the specific tradeoffs of the big three providers.
Two years ago, your data stack likely relied on a Hive Metastore (HMS) sitting behind a Thrift server. You know the pain: the "Thrift connection reset" errors at 3 AM, the sheer agony of schema mismatches, and the desperate need to perform MSCK REPAIR TABLE every time a bucket was touched. It was a brittle, monolithic bottleneck that made multi-engine access feel like a high-stakes gamble.
Today, life looks different. With the Apache Iceberg REST specification (OpenAPI), the "catalog" is no longer a database you babysit; it's an API layer that mediates between your compute engines (Spark, Trino, Flink) and your storage (S3, GCS, ADLS). You point your engine at a URL, provide an OAuth token, and the metadata just works. No more manual partition scanning. No more proprietary lock-in at the catalog layer.
The real problem
The industry calls this "decoupling," but let's be honest: it’s about avoiding the "I can't believe this engine just corrupted my data" conversation with your VP of Engineering.
The real problem with traditional metastores isn't just the protocol; it's the lack of fine-grained access control and the inability to handle concurrent writes safely. When you move to an Iceberg REST catalog, you aren't just changing a URL. You are moving to a world where the catalog defines the "source of truth" via atomic commits. If your catalog backend doesn't support the full Iceberg REST spec (version 1.0.0 and above), you’re just putting lipstick on a pig.
Photo by Yancy Min on Unsplash
Step: Deploying Apache Polaris for the pure-play approach
If you want to own your infrastructure and avoid vendor lock-in, Apache Polaris is the new gold standard. It’s built by Snowflake, but it’s open-source and specifically designed for the Iceberg REST spec.
When you deploy Polaris, you’re running a Java-based service that handles authorization and metadata management. You don’t need a backing RDBMS if you don’t want one—you can configure it to use local storage or cloud object stores. Here is the minimal configuration to get a Polaris service running in a containerized environment:
# polaris-config.yaml
server:
type: default
applicationConnectors:
- type: http
port: 8181
adminConnectors:
- type: http
port: 8182
storage:
type: s3
bucket: my-iceberg-data-bucket
region: us-east-1
auth:
type: oauth2
issuer: https://my-auth-provider.com/
The key here is the auth section. Polaris forces you to handle security properly. If you aren't using an OIDC provider, don't even bother starting. Polaris expects a token in the Authorization header, and it will reject any write request that isn't signed correctly.
Step: Leveraging Unity Catalog for the "it just works" path
If you are already in the Databricks ecosystem, Unity Catalog (UC) is the pragmatic choice. Unity Catalog has evolved to support the Iceberg REST spec, allowing you to use non-Databricks engines (like Trino or Starburst) against the same tables you manage in Databricks.
This is the "Unified Governance" play. You define your catalog and schema in the Unity Catalog UI, then pass the connection details to your external engine.
# Example Spark configuration for connecting to Unity Catalog as an Iceberg REST provider
spark.sql.catalog.uc_iceberg = org.apache.iceberg.spark.SparkCatalog
spark.sql.catalog.uc_iceberg.catalog-impl = org.apache.iceberg.rest.RESTCatalog
spark.sql.catalog.uc_iceberg.uri = https://<databricks-workspace-url>/api/2.0/unity-catalog/iceberg
spark.sql.catalog.uc_iceberg.token = <your-databricks-pat>
spark.sql.catalog.uc_iceberg.warehouse = <catalog-name>
The failure mode here is usually the token expiration. Databricks Personal Access Tokens (PATs) have a shelf life. If you’re running production pipelines, swap this for a Service Principal. If your pipeline suddenly stops writing, 9 times out of 10, the token rotated and your CI/CD environment variables didn't update.
Step: AWS Glue as the managed REST catalog
AWS Glue is the "lazy" choice, and I mean that in the best way possible. If your infrastructure is 100% AWS, using Glue as your Iceberg REST catalog backend eliminates the need to manage any servers. You get the benefit of IAM integration for free.
Glue now has a native Iceberg REST endpoint. Instead of the old GlueCatalog implementation that relied on the Hive-compatible GetTable API, you use the REST-based endpoint.
# Glue connection setup via Boto3/Environment
os.environ["ICEBERG_REST_URI"] = "https://glue.us-east-1.amazonaws.com"
os.environ["ICEBERG_CATALOG_TYPE"] = "rest"
os.environ["AWS_REGION"] = "us-east-1"
# In your Spark session:
spark.conf.set("spark.sql.catalog.glue_catalog", "org.apache.iceberg.spark.SparkCatalog")
spark.conf.set("spark.sql.catalog.glue_catalog.catalog-impl", "org.apache.iceberg.rest.RESTCatalog")
spark.conf.set("spark.sql.catalog.glue_catalog.uri", os.environ["ICEBERG_REST_URI"])
Warning: AWS Glue has hard limits on request rates. If you have a massive Flink job performing thousands of small commits per minute, you will hit the ThrottlingException wall. You’ll need to request a quota increase from AWS support before you go live, or you'll spend your first production day watching your pipelines crash.
Photo by Harshit Katiyar on Unsplash
Lessons learned from production
-
The "Namespace" trap: Iceberg REST catalogs handle namespaces differently than HMS. If you’re migrating, don't assume
my_db.my_tablein Hive will map 1:1 without testing the catalog-to-storage path. UseCALL system.migrate('db.table')carefully in a staging environment first. -
Metadata location: Always keep your metadata files (
.metadata.json) in the same bucket as your data. If you split them and your catalog backend loses connection to the metadata store, you are looking at a manual recovery process that involves editing manifest files by hand. Do not recommend. -
Engine compatibility: Not all engines support all REST features. If you are using Trino, check the Iceberg version it supports against the specific features (like branching or tagging) you want to use. A version mismatch between the catalog's API response and the engine's client library will manifest as a cryptic
404 Not Foundthat actually means "I don't understand the JSON schema you just sent me." -
Cost of observability: Unlike the old HMS which was just a MySQL database you could query, these REST catalogs are black boxes. You need to log the API requests. If you aren't logging the
POST /v1/catalogs/{catalog}/namespaces/{namespace}/tablescalls, you have zero visibility into what's failing when a write job dies.
Conclusion
The shift to Iceberg REST is the most significant improvement in data infrastructure since the invention of S3. We are finally moving away from the "database as a filesystem" hack that plagued the Hadoop era.
Whether you go with the self-hosted flexibility of Polaris, the unified governance of Unity Catalog, or the managed simplicity of AWS Glue, pick one and move your write-heavy workloads over. Your future self, sleeping soundly at 3 AM instead of debugging metastore corruption, will thank you.
Try it: Spin up a local instance of Polaris using Docker, hook it up to a local S3-compatible store like MinIO, and run a Spark job to write your first table. If you can move from a local file-based catalog to a REST-based one in under an hour, you're ready for the big leagues.
Tags: #iceberg #data #architecture #backend
Top comments (0)