DEV Community

#dataengineering

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Your Scraper Collected 50 Rows. There Were 4,000.

Your Scraper Collected 50 Rows. There Were 4,000.

Comments
7 min read
Deeper into Dataform 3: Auditing Dataform

Deeper into Dataform 3: Auditing Dataform

Comments
2 min read
How I Broke Down My ETL Pipeline Project Into Smaller Engineering Exercises

How I Broke Down My ETL Pipeline Project Into Smaller Engineering Exercises

Comments
2 min read
Apache Data Lakehouse Weekly: July 1 to July 8, 2026

Apache Data Lakehouse Weekly: July 1 to July 8, 2026

2
Comments 1
24 min read
I built a data-contract validator in pure Python (no pandas, no PyYAML) and it caught a 30% revenue ghost

I built a data-contract validator in pure Python (no pandas, no PyYAML) and it caught a 30% revenue ghost

Comments
6 min read
A practical pipeline for turning messy business documents into spreadsheets

A practical pipeline for turning messy business documents into spreadsheets

Comments
2 min read
gold 숫자가 이상할 때 source_hash, quality, lineage로 원인 좁히기

gold 숫자가 이상할 때 source_hash, quality, lineage로 원인 좁히기

Comments
2 min read
A Fair Coin Isn't Enough: When a Perfectly Randomized Experiment Is Impossible to Analyze

A Fair Coin Isn't Enough: When a Perfectly Randomized Experiment Is Impossible to Analyze

1
Comments 2
6 min read
Designing an API-First Value-Based Care Analytics Stack for MA Payers

Designing an API-First Value-Based Care Analytics Stack for MA Payers

1
Comments 1
3 min read
How do I answer "what did my data look like last month" in Postgres?

How do I answer "what did my data look like last month" in Postgres?

2
Comments 2
4 min read
Day 9 of 100 Days of ClickHouse®: Mastering Data Aggregation from GROUP BY to CUBE

Day 9 of 100 Days of ClickHouse®: Mastering Data Aggregation from GROUP BY to CUBE

1
Comments
4 min read
The Search Engine Renaissance: How Apache Lucene and Elasticsearch Are Reclaiming the AI-Native Future

The Search Engine Renaissance: How Apache Lucene and Elasticsearch Are Reclaiming the AI-Native Future

1
Comments
8 min read
If the warehouse already has the data, why are we copying it elsewhere?

If the warehouse already has the data, why are we copying it elsewhere?

Comments
5 min read
bronze, silver, and gold standard data vault 2.0

bronze, silver, and gold standard data vault 2.0

1
Comments
8 min read
The State of Apache Polaris in July 2026: From Incubating Catalog to the Governance Layer of the Open Lakehouse

The State of Apache Polaris in July 2026: From Incubating Catalog to the Governance Layer of the Open Lakehouse

Comments 1
22 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.