DEV Community

#dataengineering

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Apache Iceberg Query Optimization: Production Guide

Apache Iceberg Query Optimization: Production Guide

Comments
40 min read
Bulk YouTube transcript extraction for AI pipelines: what breaks at scale

Bulk YouTube transcript extraction for AI pipelines: what breaks at scale

Comments
4 min read
Apache Iceberg Governance: Access Control, Policies, and Audit for Open Lakehouses

Apache Iceberg Governance: Access Control, Policies, and Audit for Open Lakehouses

Comments
19 min read
Price matching is a bad default: model the pricing decision instead

Price matching is a bad default: model the pricing decision instead

Comments
5 min read
What a Delta table actually is: Parquet files plus a transaction log

What a Delta table actually is: Parquet files plus a transaction log

Comments
3 min read
We nearly charged our own buyers twice for rows they'd already paid for

We nearly charged our own buyers twice for rows they'd already paid for

Comments
4 min read
Data Engineering: Pahlawan Tanpa Tanda Jasa di Balik Era Big Data - 22:09

Data Engineering: Pahlawan Tanpa Tanda Jasa di Balik Era Big Data - 22:09

Comments
2 min read
What Splink actually runs on DuckDB when it scores entity pairs

What Splink actually runs on DuckDB when it scores entity pairs

3
Comments
2 min read
Masa Depan Manajemen Data: Mengenal Konsep Data Mesh yang Revolusioner

Masa Depan Manajemen Data: Mengenal Konsep Data Mesh yang Revolusioner

Comments
2 min read
Databricks classic vs serverless compute: the limitation list decides it

Databricks classic vs serverless compute: the limitation list decides it

Comments
2 min read
Spark Is a Smart Engine. So Why Doesn't It Cache Automatically?

Spark Is a Smart Engine. So Why Doesn't It Cache Automatically?

Comments
3 min read
We published our fundraising-prediction model and its misses. Here's what 219 rounds taught us.

We published our fundraising-prediction model and its misses. Here's what 219 rounds taught us.

Comments 1
3 min read
We timed HTTP-only scraping against a headless browser on the same page. It wasn't close.

We timed HTTP-only scraping against a headless browser on the same page. It wasn't close.

Comments
4 min read
Clean a million rows with Hydra ETL — no database, no Docker

Clean a million rows with Hydra ETL — no database, no Docker

Comments
6 min read
Databricks accounts, workspaces and metastores: which layer owns what

Databricks accounts, workspaces and metastores: which layer owns what

Comments
2 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.