DEV Community

Gowtham Potureddi profile picture

Gowtham Potureddi

404 bio not found

Joined Joined on 
Prophecy: Visual, Git-Backed Low-Code Spark & SQL Pipelines for the Enterprise

Prophecy: Visual, Git-Backed Low-Code Spark & SQL Pipelines for the Enterprise

Comments
26 min read
Bruin: SQL + Python Pipelines in One Framework With Built-In Quality Checks

Bruin: SQL + Python Pipelines in One Framework With Built-In Quality Checks

Comments
27 min read
A Local Lakehouse on Your Laptop: DuckDB + Iceberg + Trino for Zero-Cloud Dev

A Local Lakehouse on Your Laptop: DuckDB + Iceberg + Trino for Zero-Cloud Dev

Comments
28 min read
Data Residency & Sovereignty: Multi-Region Architectures for Compliant Data Platforms

Data Residency & Sovereignty: Multi-Region Architectures for Compliant Data Platforms

Comments
32 min read
GDPR vs CCPA vs India DPDP: Building Multi-Jurisdiction Privacy Pipelines

GDPR vs CCPA vs India DPDP: Building Multi-Jurisdiction Privacy Pipelines

Comments
29 min read
SOC 2 for Data Engineering: Controls, Audit Logs & Evidence-Collection Pipelines

SOC 2 for Data Engineering: Controls, Audit Logs & Evidence-Collection Pipelines

Comments
29 min read
Marimo: The Reactive Python Notebook for Reproducible Data Work

Marimo: The Reactive Python Notebook for Reproducible Data Work

Comments
26 min read
Papermill & Notebook Pipelines: Parametrized, Scheduled, Version-Controlled Notebooks

Papermill & Notebook Pipelines: Parametrized, Scheduled, Version-Controlled Notebooks

Comments
25 min read
Streamlit for Data Engineers: Internal Data Apps & Pipeline Dashboards

Streamlit for Data Engineers: Internal Data Apps & Pipeline Dashboards

Comments
29 min read
Factless Fact Tables: Modeling Events, Coverage & Eligibility Without Measures

Factless Fact Tables: Modeling Events, Coverage & Eligibility Without Measures

Comments
28 min read
Bridge Tables & Many-to-Many Dimensions: Modeling Hierarchies and Multi-Valued Attributes

Bridge Tables & Many-to-Many Dimensions: Modeling Hierarchies and Multi-Valued Attributes

Comments
31 min read
Special Dimensions: Junk, Degenerate, Role-Playing & Conformed Dimensions

Special Dimensions: Junk, Degenerate, Role-Playing & Conformed Dimensions

Comments
29 min read
Fact Table Patterns: Transaction, Periodic Snapshot & Accumulating Snapshot Facts

Fact Table Patterns: Transaction, Periodic Snapshot & Accumulating Snapshot Facts

Comments
29 min read
Iceberg Table Maintenance: Compaction, Expire Snapshots, Rewrite Manifests & Orphan Files

Iceberg Table Maintenance: Compaction, Expire Snapshots, Rewrite Manifests & Orphan Files

Comments
30 min read
Apache Amoro (was Arctic): Self-Optimizing Lakehouse Management for Iceberg & Paimon

Apache Amoro (was Arctic): Self-Optimizing Lakehouse Management for Iceberg & Paimon

Comments
28 min read
Apache Paimon: The Streaming Lakehouse Table Format for Flink & Spark

Apache Paimon: The Streaming Lakehouse Table Format for Flink & Spark

Comments
29 min read
Databricks SQL Warehouses & Serverless Compute: Sizing, Photon & Cost Control

Databricks SQL Warehouses & Serverless Compute: Sizing, Photon & Cost Control

Comments
28 min read
Google Cloud Composer: Managed Airflow on GCP — Sizing, Tuning & Gotchas

Google Cloud Composer: Managed Airflow on GCP — Sizing, Tuning & Gotchas

Comments
29 min read
Google Cloud Dataproc & Dataproc Serverless: Managed Spark on GCP

Google Cloud Dataproc & Dataproc Serverless: Managed Spark on GCP

Comments
27 min read
Amazon EMR Deep Dive: Cluster Types, Spot Fleets, EMR Serverless & Iceberg

Amazon EMR Deep Dive: Cluster Types, Spot Fleets, EMR Serverless & Iceberg

Comments
30 min read
Azure Databricks vs Microsoft Fabric: Choosing Your Azure Lakehouse in 2026

Azure Databricks vs Microsoft Fabric: Choosing Your Azure Lakehouse in 2026

Comments
29 min read
ADLS Gen2 for Data Engineers: Hierarchical Namespace, POSIX ACLs & Performance

ADLS Gen2 for Data Engineers: Hierarchical Namespace, POSIX ACLs & Performance

Comments
29 min read
Azure Synapse Analytics Deep Dive: Dedicated vs Serverless SQL Pools & Spark Pools

Azure Synapse Analytics Deep Dive: Dedicated vs Serverless SQL Pools & Spark Pools

Comments
28 min read
Azure Data Factory Deep Dive: Mapping Data Flows, Triggers, Integration Runtimes & CI/CD

Azure Data Factory Deep Dive: Mapping Data Flows, Triggers, Integration Runtimes & CI/CD

Comments
27 min read
SQS, SNS & EventBridge: Simple Queues and Fan-Out in AWS Data Pipelines

SQS, SNS & EventBridge: Simple Queues and Fan-Out in AWS Data Pipelines

Comments
30 min read
NATS & JetStream: Lightweight Messaging for Edge & Real-Time Pipelines

NATS & JetStream: Lightweight Messaging for Edge & Real-Time Pipelines

Comments
30 min read
RabbitMQ vs Kafka for Data Engineering: Queues vs Logs, When Each Wins

RabbitMQ vs Kafka for Data Engineering: Queues vs Logs, When Each Wins

Comments
29 min read
Amazon MSK & MSK Serverless: Managed Kafka Without the Ops

Amazon MSK & MSK Serverless: Managed Kafka Without the Ops

Comments
32 min read
PyAirbyte: Running Airbyte Connectors as a Python Library

PyAirbyte: Running Airbyte Connectors as a Python Library

Comments
26 min read
Sling: CLI-First Database-to-Database & File Replication for Data Teams

Sling: CLI-First Database-to-Database & File Replication for Data Teams

Comments
27 min read
dlt (data load tool) for Data Engineers: Schema Inference, Incremental Loads & Load Modes

dlt (data load tool) for Data Engineers: Schema Inference, Incremental Loads & Load Modes

Comments
24 min read
Makefiles & Taskfiles for Reproducible Data Workflows

Makefiles & Taskfiles for Reproducible Data Workflows

Comments
62 min read
Structured Logging for Data Pipelines: JSON Logs & Correlation IDs

Structured Logging for Data Pipelines: JSON Logs & Correlation IDs

Comments
63 min read
uv: The Fast Python Package & Project Manager for Data Teams

uv: The Fast Python Package & Project Manager for Data Teams

Comments
64 min read
Data Skew Explained: Why One Task Runs Forever (and How to Fix It)

Data Skew Explained: Why One Task Runs Forever (and How to Fix It)

Comments
65 min read
Big-Data Partitioning Strategies: Range, Hash & List

Big-Data Partitioning Strategies: Range, Hash & List

Comments
69 min read
Row vs Columnar Storage: Why It Matters

Row vs Columnar Storage: Why It Matters

Comments
70 min read
The Modern Data Stack in 2026, Explained Simply

The Modern Data Stack in 2026, Explained Simply

Comments
61 min read
Dremio & Query Acceleration: Reflections on the Open Lakehouse

Dremio & Query Acceleration: Reflections on the Open Lakehouse

Comments
64 min read
Apache Iceberg v3: Deletion Vectors, Row Lineage & Binary Types

Apache Iceberg v3: Deletion Vectors, Row Lineage & Binary Types

Comments
61 min read
Zero-Copy Cloning: Instant Dev & Test Data in Snowflake & Databricks

Zero-Copy Cloning: Instant Dev & Test Data in Snowflake & Databricks

Comments
62 min read
Zero-ETL Explained: Aurora, DynamoDB & Salesforce Warehouse

Zero-ETL Explained: Aurora, DynamoDB & Salesforce Warehouse

Comments
61 min read
Azure Event Hubs & Functions: Event-Driven Data on Azure

Azure Event Hubs & Functions: Event-Driven Data on Azure

Comments
62 min read
Google Pub/Sub & Dataflow: Streaming Pipelines on GCP

Google Pub/Sub & Dataflow: Streaming Pipelines on GCP

Comments
59 min read
Amazon Kinesis Deep Dive: Data Streams, Firehose & Analytics

Amazon Kinesis Deep Dive: Data Streams, Firehose & Analytics

Comments
66 min read
AWS Step Functions: Serverless Orchestration for Data

AWS Step Functions: Serverless Orchestration for Data

Comments
60 min read
AWS Lambda for ETL: Event-Driven Data Pipelines

AWS Lambda for ETL: Event-Driven Data Pipelines

Comments
68 min read
Redis for Data Engineers: Beyond Caching

Redis for Data Engineers: Beyond Caching

Comments
62 min read
MongoDB for Data Engineers: Aggregation Pipeline & $lookup

MongoDB for Data Engineers: Aggregation Pipeline & $lookup

Comments
54 min read
Cassandra & ScyllaDB: Wide-Column Data Modeling Done Right

Cassandra & ScyllaDB: Wide-Column Data Modeling Done Right

Comments
59 min read
DynamoDB for Data Engineers: Single-Table Design, Streams & S3 Export

DynamoDB for Data Engineers: Single-Table Design, Streams & S3 Export

Comments
57 min read
Sketches at Scale: Count-Min, t-digest & Approximate Quantiles

Sketches at Scale: Count-Min, t-digest & Approximate Quantiles

Comments
70 min read
Consistent Hashing: How Distributed Systems Partition Data

Consistent Hashing: How Distributed Systems Partition Data

Comments
63 min read
HyperLogLog: Count Billions of Uniques in Kilobytes

HyperLogLog: Count Billions of Uniques in Kilobytes

Comments
66 min read
Bloom Filters for Data Engineers: Cheap Membership Tests

Bloom Filters for Data Engineers: Cheap Membership Tests

Comments
55 min read
Retries, Timeouts & Circuit Breakers for Data Pipelines

Retries, Timeouts & Circuit Breakers for Data Pipelines

Comments
65 min read
Dead Letter Queues: Never Lose a Bad Record

Dead Letter Queues: Never Lose a Bad Record

Comments
60 min read
Handling Late & Out-of-Order Data

Handling Late & Out-of-Order Data

Comments
73 min read
Backfilling Data Without Breaking Production

Backfilling Data Without Breaking Production

Comments
72 min read
Idempotent Data Pipelines: Safe Retries Without Duplicates

Idempotent Data Pipelines: Safe Retries Without Duplicates

Comments
60 min read
loading...