DEV Community

Cover image for Death to the Hive Metastore: Why Iceberg REST Catalogs Are Your New Reality
Aniket Abhishek Soni
Aniket Abhishek Soni

Posted on

Death to the Hive Metastore: Why Iceberg REST Catalogs Are Your New Reality

Why I chose this topic: Three years ago, I spent an entire weekend debugging a corrupted Hive Metastore (HMS) that locked up our daily reconciliation pipeline. I'm writing this because the industry is finally moving toward decoupled, REST-based cataloging, and I want you to avoid the migration headaches I dealt with by learning the specific tradeoffs of the big three providers.

Two years ago, your data stack likely relied on a Hive Metastore (HMS) sitting behind a Thrift server. You know the pain: the "Thrift connection reset" errors at 3 AM, the sheer agony of schema mismatches, and the desperate need to perform MSCK REPAIR TABLE every time a bucket was touched. It was a brittle, monolithic bottleneck that made multi-engine access feel like a high-stakes gamble.

Today, life looks different. With the Apache Iceberg REST specification (OpenAPI), the "catalog" is no longer a database you babysit; it's an API layer that mediates between your compute engines (Spark, Trino, Flink) and your storage (S3, GCS, ADLS). You point your engine at a URL, provide an OAuth token, and the metadata just works. No more manual partition scanning. No more proprietary lock-in at the catalog layer.

The real problem

The industry calls this "decoupling," but let's be honest: it’s about avoiding the "I can't believe this engine just corrupted my data" conversation with your VP of Engineering.

The real problem with traditional metastores isn't just the protocol; it's the lack of fine-grained access control and the inability to handle concurrent writes safely. When you move to an Iceberg REST catalog, you aren't just changing a URL. You are moving to a world where the catalog defines the "source of truth" via atomic commits. If your catalog backend doesn't support the full Iceberg REST spec (version 1.0.0 and above), you’re just putting lipstick on a pig.

Photo by Yancy Min on Unsplash
Photo by Yancy Min on Unsplash

Step: Deploying Apache Polaris for the pure-play approach

If you want to own your infrastructure and avoid vendor lock-in, Apache Polaris is the new gold standard. It’s built by Snowflake, but it’s open-source and specifically designed for the Iceberg REST spec.

When you deploy Polaris, you’re running a Java-based service that handles authorization and metadata management. You don’t need a backing RDBMS if you don’t want one—you can configure it to use local storage or cloud object stores. Here is the minimal configuration to get a Polaris service running in a containerized environment:

# polaris-config.yaml
server:
  type: default
  applicationConnectors:
    - type: http
      port: 8181
  adminConnectors:
    - type: http
      port: 8182

storage:
  type: s3
  bucket: my-iceberg-data-bucket
  region: us-east-1

auth:
  type: oauth2
  issuer: https://my-auth-provider.com/
Enter fullscreen mode Exit fullscreen mode

The key here is the auth section. Polaris forces you to handle security properly. If you aren't using an OIDC provider, don't even bother starting. Polaris expects a token in the Authorization header, and it will reject any write request that isn't signed correctly.

Step: Leveraging Unity Catalog for the "it just works" path

If you are already in the Databricks ecosystem, Unity Catalog (UC) is the pragmatic choice. Unity Catalog has evolved to support the Iceberg REST spec, allowing you to use non-Databricks engines (like Trino or Starburst) against the same tables you manage in Databricks.

This is the "Unified Governance" play. You define your catalog and schema in the Unity Catalog UI, then pass the connection details to your external engine.

# Example Spark configuration for connecting to Unity Catalog as an Iceberg REST provider
spark.sql.catalog.uc_iceberg = org.apache.iceberg.spark.SparkCatalog
spark.sql.catalog.uc_iceberg.catalog-impl = org.apache.iceberg.rest.RESTCatalog
spark.sql.catalog.uc_iceberg.uri = https://<databricks-workspace-url>/api/2.0/unity-catalog/iceberg
spark.sql.catalog.uc_iceberg.token = <your-databricks-pat>
spark.sql.catalog.uc_iceberg.warehouse = <catalog-name>
Enter fullscreen mode Exit fullscreen mode

The failure mode here is usually the token expiration. Databricks Personal Access Tokens (PATs) have a shelf life. If you’re running production pipelines, swap this for a Service Principal. If your pipeline suddenly stops writing, 9 times out of 10, the token rotated and your CI/CD environment variables didn't update.

Step: AWS Glue as the managed REST catalog

AWS Glue is the "lazy" choice, and I mean that in the best way possible. If your infrastructure is 100% AWS, using Glue as your Iceberg REST catalog backend eliminates the need to manage any servers. You get the benefit of IAM integration for free.

Glue now has a native Iceberg REST endpoint. Instead of the old GlueCatalog implementation that relied on the Hive-compatible GetTable API, you use the REST-based endpoint.

# Glue connection setup via Boto3/Environment
os.environ["ICEBERG_REST_URI"] = "https://glue.us-east-1.amazonaws.com"
os.environ["ICEBERG_CATALOG_TYPE"] = "rest"
os.environ["AWS_REGION"] = "us-east-1"

# In your Spark session:
spark.conf.set("spark.sql.catalog.glue_catalog", "org.apache.iceberg.spark.SparkCatalog")
spark.conf.set("spark.sql.catalog.glue_catalog.catalog-impl", "org.apache.iceberg.rest.RESTCatalog")
spark.conf.set("spark.sql.catalog.glue_catalog.uri", os.environ["ICEBERG_REST_URI"])
Enter fullscreen mode Exit fullscreen mode

Warning: AWS Glue has hard limits on request rates. If you have a massive Flink job performing thousands of small commits per minute, you will hit the ThrottlingException wall. You’ll need to request a quota increase from AWS support before you go live, or you'll spend your first production day watching your pipelines crash.

Photo by Harshit Katiyar on Unsplash
Photo by Harshit Katiyar on Unsplash

Lessons learned from production

  • The "Namespace" trap: Iceberg REST catalogs handle namespaces differently than HMS. If you’re migrating, don't assume my_db.my_table in Hive will map 1:1 without testing the catalog-to-storage path. Use CALL system.migrate('db.table') carefully in a staging environment first.
  • Metadata location: Always keep your metadata files (.metadata.json) in the same bucket as your data. If you split them and your catalog backend loses connection to the metadata store, you are looking at a manual recovery process that involves editing manifest files by hand. Do not recommend.
  • Engine compatibility: Not all engines support all REST features. If you are using Trino, check the Iceberg version it supports against the specific features (like branching or tagging) you want to use. A version mismatch between the catalog's API response and the engine's client library will manifest as a cryptic 404 Not Found that actually means "I don't understand the JSON schema you just sent me."
  • Cost of observability: Unlike the old HMS which was just a MySQL database you could query, these REST catalogs are black boxes. You need to log the API requests. If you aren't logging the POST /v1/catalogs/{catalog}/namespaces/{namespace}/tables calls, you have zero visibility into what's failing when a write job dies.

Conclusion

The shift to Iceberg REST is the most significant improvement in data infrastructure since the invention of S3. We are finally moving away from the "database as a filesystem" hack that plagued the Hadoop era.

Whether you go with the self-hosted flexibility of Polaris, the unified governance of Unity Catalog, or the managed simplicity of AWS Glue, pick one and move your write-heavy workloads over. Your future self, sleeping soundly at 3 AM instead of debugging metastore corruption, will thank you.

Try it: Spin up a local instance of Polaris using Docker, hook it up to a local S3-compatible store like MinIO, and run a Spark job to write your first table. If you can move from a local file-based catalog to a REST-based one in under an hour, you're ready for the big leagues.


Tags: #iceberg #data #architecture #backend

Cover photo by Kirill Sh on Unsplash.

Top comments (0)