DEV Community

Cover image for Cutting AWS Storage Costs: S3 Tables vs. S3 Express One Zone
Aniket Abhishek Soni
Aniket Abhishek Soni

Posted on

Cutting AWS Storage Costs: S3 Tables vs. S3 Express One Zone

Last year, my team was fighting the "small file problem" in a standard S3 bucket. We were running massive Glue jobs that spent 40% of their execution time just doing LIST requests to aggregate millions of tiny Parquet files. We had s3:ListBucket throttled, our EMR clusters were idling during metadata scanning, and our monthly storage bill looked like a cry for help.

Today, we’ve shifted our hot-path workloads to S3 Express One Zone and our long-term cataloged tables to S3 Tables. The result isn't just "better performance." We went from 30-minute metadata listing times to sub-second responses. More importantly, we stopped paying for the privilege of our own inefficiency.

Most engineers think of S3 as a simple key-value store. It’s not. It’s a distributed system that behaves fundamentally differently depending on the storage class you pick. If you treat S3 Express One Zone like a standard Standard bucket, you’re setting fire to your AWS credits.

How it actually works

To understand why S3 Tables and S3 Express One Zone change the game, you have to look at the metadata tax. In standard S3, LIST operations are expensive and eventually consistent. When you have a table with 50 million files, your query engine (Trino, Spark, or Athena) has to crawl that bucket. That’s a network-heavy, latency-riddled nightmare.

S3 Express One Zone (the S3EX storage class) changes the architecture by co-locating compute and storage within a single AZ. It uses a custom directory-bucket structure. Instead of the flat namespace of standard S3, you are using a specialized API that provides single-digit millisecond latency.

The code change is trivial but massive in impact. You move from the standard s3:// URI to s3express:// (or use the bucket name with the --onezone suffix).

When you configure your Spark session, you aren't just changing a prefix. You are bypassing the standard S3 global endpoint.

// The old way: slow, throttled, expensive
val df = spark.read.parquet("s3a://my-bucket/data/logs/")

// The new way: S3 Express One Zone
val df = spark.read.parquet("s3express://my-bucket--use1-az1--x-s3/data/logs/")
Enter fullscreen mode Exit fullscreen mode

S3 Tables, on the other hand, is the management layer. It’s essentially a managed Apache Iceberg engine built directly into S3. It abstracts the file management. You no longer worry about file compaction jobs because S3 Tables handles the vacuuming and snapshot management for you. You interact with it via the s3tables API, and it handles the underlying metadata manifest files. You effectively trade control for reduced operational toil.

Photo by 灿雄 邱 on Unsplash
Photo by 灿雄 邱 on Unsplash

The tradeoffs nobody mentions

Let’s be honest: AWS doesn't give you these performance gains for free. There is a "Single Zone" tax, and it’s not just about durability.

With S3 Express One Zone, you lose regional redundancy. If your chosen AZ goes dark, your data is unavailable. In healthcare or financial services, that’s a non-starter for PII or critical audit logs unless you have an automated cross-region replication strategy in place—which, by the way, eats up all the cost savings you just gained.

Then there is the API cost. S3 Express One Zone charges per request, but those costs are bundled differently. If you are doing infrequent, massive sequential reads, standard S3 is still cheaper. If you are doing high-frequency, small-file random access (think real-time features engineering or feature stores), S3 Express One Zone is significantly cheaper because you aren't paying the overhead of the standard S3 request tax for millions of tiny metadata operations.

S3 Tables has its own headache: lock-in. Once you start using the S3 Tables managed format, moving that data out to a different provider or an on-prem cluster is significantly harder than moving standard Parquet files. You are tied to the AWS-managed Iceberg implementation. If you love open-source portability, this is a bitter pill to swallow.

Also, be warned about the "One Zone" prefix. If you are running your EMR or Glue jobs in us-east-1a and your S3 Express bucket is in us-east-1b, you are going to eat data transfer costs that will destroy your budget. You must ensure your compute and your S3EX bucket are pinned to the exact same AZ. If you don't have a strict infrastructure-as-code (IaC) policy enforcing this, your cloud bill will fluctuate wildly based on which AZ your spot instances happen to land in.

Photo by Agil Saputro on Unsplash
Photo by Agil Saputro on Unsplash

When to reach for it (and when not to)

Use S3 Express One Zone when your bottleneck is IOPS, not throughput. If your Spark jobs are stuck in "waiting for metadata" or "listing directory," this is the fix. It’s perfect for training machine learning models where you need to stream millions of small image files or feature vectors into a GPU cluster.

Do not use it for your primary "cold" data lake. Your quarterly historical reports don't need sub-millisecond access. Keeping those in S3EX is just bad engineering. Keep your cold data in S3 Standard or S3 Intelligent-Tiering.

Reach for S3 Tables when your team is drowning in Iceberg maintenance. If you find yourself writing custom scripts to run VACUUM and OPTIMIZE on your tables, S3 Tables is your escape hatch. It turns the "data engineer as a janitor" role into something more productive.

However, avoid S3 Tables if you require strict, fine-grained control over your partition evolution or if you have complex custom manifest requirements that the AWS managed service doesn't support yet. If you have a highly customized, non-standard Iceberg implementation, the migration process will be a weekend-ruining experience.

Conclusion

The era of just "dumping everything in a bucket" is over. We’ve been using S3 as a giant, undifferentiated blob store for a decade, and we’re paying for it with massive latency and wasted compute cycles.

By separating our storage into tiered access—S3 Tables for the managed, queryable lakehouse and S3 Express One Zone for the high-frequency IO path—we’ve actually started to see our storage costs move in the right direction. It requires more thoughtful architecture and a stricter handle on your AZ affinity, but that’s the job. Stop treating your data lake like a digital landfill and start treating it like a distributed database. Your infrastructure costs will thank you.

Cover photo by Albert Stoynov on Unsplash.

Top comments (0)