DEV Community

Apache SeaTunnel
Apache SeaTunnel

Posted on

Countdown: 1 Day to Community Over Code Asia 2026 — Apache SeaTunnel Brings 5 Technical Sessions Starting August 8

Only 1 day to go! Apache SeaTunnel sessions at Community Over Code Asia 2026 will kick off on August 8 (China Standard Time).

As one of the Apache Software Foundation’s key technology events in Asia, Community Over Code Asia 2026 will bring together open source contributors, developers, and technology enthusiasts from around the world to explore the latest practices and future trends in open source technology.

Starting from August 8 (China Standard Time), Apache SeaTunnel will bring five technical sessions across two major tracks — Data Lake & Data Warehouse and DataOps — covering key topics including data integration, real-time synchronization, AI-driven data pipelines, and the evolution of modern data infrastructure.

Before the event begins, let’s take a closer look at the Apache SeaTunnel sessions and explore the technical insights behind each talk. We warmly invite you to join the event, connect with the Apache SeaTunnel community, and be part of this global open source gathering.

See you on August 8!

Data Lake & Data Warehouse Track

Data lakes and data warehouses are essential solutions for storing and managing data. They play a critical role in data management, analytics, and decision-making.

Within the Apache Software Foundation, many projects focus on data lakes and data warehouses, including Apache Hive, Apache Hudi, Apache Iceberg, Apache Paimon, Apache Cassandra, and Apache HBase.

In this track, attendees will learn about the latest developments in data lake and warehouse technologies, production best practices from companies, and the future roadmaps of these projects.

Track Producers

Lidong Dai

Lidong Dai is CTO of WhaleOps, Apache Incubator Mentor, Apache DolphinScheduler PMC Member, and Apache SeaTunnel PMC Member.

With 16 years of experience in data technologies, he focuses on AI-ready heterogeneous data integration, data processing orchestration, and data governance.

Based on the Apache SeaTunnel and Apache DolphinScheduler ecosystems, he leads the development of WhaleStudio, a commercial data platform solution that has served customers across industries including finance, automotive, gaming, manufacturing, retail, and internet services.

Shaofeng Shi

Shaofeng Shi is a member of the Apache Software Foundation and a Mentor of Apache Gravitino, Apache Gluten, Apache HoraeDB, and other projects.

He focuses on big data analytics and cloud computing technologies. Previously, he worked as a Senior Engineer in eBay’s Global Analytics Infrastructure team and as a Software Architect in IBM’s Cloud Computing division.

Zongtang Hu

Zongtang Hu is a technical expert in middleware and big data at China Mobile Cloud Center, where he leads the middleware and big data team.

He has more than eight years of experience in message middleware kernel development and architecture design. He participated in the development of multiple major middleware products, including China Mobile Cloud RocketMQ, MQTT, and Kafka, contributing to their core architecture and implementation from the ground up.

As a technical speaker, he has shared insights at ApacheCon Asia 2022/2023/2024, Apache RocketMQ Summit/Meetup, and Cloud Native Service conferences.

He has extensive open source community experience and serves as a Maintainer/Committer in communities including Apache RocketMQ, Nacos, openEuler message-middleware SIG, and OpenMessaging.

He was recognized as an Outstanding Contributor to Cloud Computing Open Source Standards by the China Academy of Information and Communications Technology (CAICT) in 2023, named an OSCAR Open Source Leader by CAICT in 2024, and received multiple open source community recognitions.

Huaxin Gao

Huaxin Gao is a Software Engineer at Snowflake, an Apache Spark Committer and PMC Member.

She is also a Committer of Apache Iceberg and Apache DataFusion Comet. Her contributions cover query engines, table formats, and distributed data systems.


Within this track, Apache SeaTunnel will bring one technical sessions. Let’s take a closer look.

Session: From Data Ingestion to Data Lake: Building a Modern Lakehouse with Apache SeaTunnel

Time: August 8, 14:30–15:00 (China Standard Time)

Session Overview

Building a data lake is no longer just about choosing Iceberg, Hudi, or Paimon. In real-world systems, the biggest challenge often lies one step earlier: how data reliably, efficiently, and continuously enters the lake. In this session, we will explore how Apache SeaTunnel serves as a unified data ingestion and integration layer for modern data lake architectures. Starting from common pain points—multi-source data, Batch + CDC coexistence, schema evolution, and operational complexity—we will walk through how SeaTunnel simplifies data movement into data lakes and lakehouse systems. Through real production scenarios, you will see how SeaTunnel connects transactional databases, message queues, and file systems into Iceberg- or lakehouse-based storage, enabling scalable, maintainable, and evolvable data platforms. The talk focuses on practical architecture decisions, not vendor-specific solutions.

The session will cover:

  • Typical data lake architecture evolution and common pitfalls
  • The role of data integration in lake and lakehouse systems
  • Apache SeaTunnel architecture and design principles
  • End-to-end ingestion examples: databases, CDC, and streaming data into data lakes
  • Operational considerations and best practices
  • Roadmap of SeaTunnel in the data lake ecosystem

Speaker

Lidong Dai | Co-founder of WhaleOps Technology

Apache Incubator Mentor, Apache DolphinScheduler PMC Member, and Apache SeaTunnel PMC Member.


Beyond the Data Lake & Data Warehouse track, Apache SeaTunnel will also bring four technical sessions to the DataOps track this year. Each session is packed with practical insights and cutting-edge exploration, making this one of the most anticipated parts of the event.

DataOps Track

This track focuses on some of the most innovative and forward-looking projects in the Apache ecosystem.

It brings together leading experts and contributors from projects including Apache DolphinScheduler, Apache Airflow, Apache SeaTunnel, Apache Flume, Apache Sqoop, Apache Griffin, Apache Atlas, and other DataOps-related communities to explore the latest advancements in data operations, automation, and orchestration.

Whether you are an experienced data professional or just getting started in the field, this track provides valuable insights across a wide range of topics, including data pipelines, ETL, orchestration, data quality, metadata management, and more.

Join us at Community Over Code Asia 2026 to explore the evolving world of DataOps with the Apache community.

Track Producers

Wei Guo

Community Over Code Asia 2026

CEO of WhaleOps, Apache Member, Apache Incubator Mentor

Wei Guo is the CEO of WhaleOps, a member of the Apache Software Foundation, an Apache Incubator Mentor, and currently serves as a member of the Open Source Technology Committee of the China Communications Society and Deputy Director of the Intelligent Application Services Branch of the China Software Industry Association.

He is also Vice President of the Global SME Entrepreneurship Union, President of the Beijing Chapter of TGO Kunpeng Club, and a recognized digital technology leader.

He serves as a member of the Apache DolphinScheduler PMC and an Apache SeaTunnel Mentor, and is the founder of the ClickHouse Chinese Community.

Wei Guo graduated from Peking University and has more than 20 years of experience in the big data field. He previously worked as a Senior Architect at IBM and Teradata, CTO of eBay’s China business, General Manager of Wanda E-commerce Data Department, and Director of Big Data at Lenovo Research Institute.

He has made significant contributions to research and innovation in emerging big data technologies.

Lifeng Nie

Community Over Code Asia 2026

COO of WhaleOps, Apache SeaTunnel PMC Member & Apache DolphinScheduler Committer

Lifeng Nie is COO of WhaleOps, an Apache SeaTunnel PMC Member, Apache DolphinScheduler Committer, and was recognized as one of the “33 Open Source Pioneers of China 2023.”

He also leads the volunteer team of the ClickHouse Chinese Community and actively contributes to open source ecosystem development.


1. Session: From Natural Language to Reliable Data Pipelines: Building an AI-Powered CLI for Apache SeaTunnel

Time: August 9, 13:30–14:00 (China Standard Time)

Session Overview

Modern DataOps is not only about moving data faster, but also about making pipeline development easier, safer, and more accessible to engineers and data teams.

In this talk, I will introduce seatunnel-cli, a new Python-based CLI for Apache SeaTunnel that generates HOCON pipeline configurations directly from natural language descriptions in both English and Chinese.

The tool is designed as an AI-powered multi-agent workflow: Planner, Config Generator, Validator, and Auto-fix. It combines a three-tier knowledge base, a connector catalog automatically generated from SeaTunnel Java source code, dry-run validation, and iterative repair to help users produce runnable pipeline configurations with less manual effort. It also supports multiple LLM providers, persistent session memory, interactive exploration, and single-shot scripting usage.I will share the technical design behind the CLI, including how we extract and resolve connector metadata at scale, how validation and auto-fix loops improve configuration quality, and how this approach can reduce the barrier to building SeaTunnel jobs in real-world DataOps scenarios.

The session will also cover practical lessons from testing across multilingual inputs, multiple model providers, and broken-config recovery cases.

Speaker

Xin Zhang | Solution Architect, Amazon Web Services

Xin Zhang is an AWS Solutions Architect, responsible for solution consulting and design based on the AWS Cloud platform. He has a rich experience in R&D and architecture practice in the fields of system architecture, data warehousing, and real-time computing.

2. Session: Building AI's Data Artery: Architecture and Practices of Unified Multimodal Data Pipelines

Time: August 9, 14:00–14:30 (China Standard Time)

Session Overview

In the GenAI era, the massive flow of multimodal data demands a robust infrastructure, yet fragmented data pipelines have become a critical bottleneck for enterprises. At Tongcheng Travel, we historically operated 4 disjointed data pipeline services (Offline Sync, Real-time Lake Ingestion, legacy Sqoop, and a standalone SeaTunnel service). This fragmentation caused extremely high maintenance costs and hindered unified data governance.

This session details how we successfully architected a unified “Data Artery” through platformization. We will explore how we consolidated the data entry points and built a true “Batch-Stream Unified” foundational architecture based on Apache SeaTunnel, comprehensively supporting data flows from traditional data warehouses to modern AI scenarios.

Key Content:

  1. Breaking Data Silos: A deep dive into designing a unified multimodal data pipeline service powered by the Apache SeaTunnel engine, smoothly replacing and consolidating 4 legacy integration systems to achieve complete architectural standardization.
  2. Compute Enhancement & AI Multimodal Empowerment: Exploring how to deeply integrate SeaTunnel’s Transform mechanism with real-time stream processing capabilities to efficiently execute complex data cleaning and dynamic transformations. We will highlight hardcore support for AI workloads, including real-time parsing of unstructured data and Embedding preprocessing for LLMs.
  3. Zero-Downtime Migration & Strict Validation: Sharing enterprise-grade practices on migrating massive legacy tasks. We will detail the “Dynamic Task Conversion and Bi-directional Data Reconciliation Mechanism” we designed to ensure zero data loss and a seamless transition for the business during the underlying architecture upgrade.
  4. Future Cloud-Native Evolution: Looking ahead at the blueprint for multimodal unified data pipelines. We will discuss cloud-native containerized deployments on Kubernetes for elastic scaling, and how to deeply integrate with the LLM ecosystem to build a robust Data+AI foundation.

Speaker

Xiaochen Zhou | Data Engineer at Tongcheng Travel | Apache SeaTunnel Committer

Xiaochen Zhou is a Data Engineer at Tongcheng Travel and an active Apache SeaTunnel Committer. In his current role, he specializes in designing, building, and optimizing high-performance data pipelines. Within the open-source community, he is deeply involved in the core development and technical evolution of Apache SeaTunnel. Recently, his focus has shifted to the intersection of Data and AI, where he is dedicated to architecting unified multimodal data pipelines for the GenAI era.


3.From 'Usable' to 'Governable': ClassLoader Lifecycle Governance Practice in Apache SeaTunnel

Time: August 9, 15:15–15:45 (China Standard Time)

Session Overview

ClassLoader leaks are among the most hidden and difficult-to-diagnose runtime problems in long-lived JVM systems, and a common challenge faced by long-running JVM workloads across the Apache ecosystem. Existing approaches largely remain at the “post-mortem investigation” stage: monitoring tools and heap dumps can tell developers “which ClassLoaders are still alive,” but struggle to clearly explain “why they cannot be reclaimed,” let alone translate governance intent into verifiable, reproducible engineering practices.

Through deep analysis of Apache SeaTunnel’s classloading mechanism, we identified a widely overlooked blind spot: a large number of seemingly correct lifecycle implementations never establish explicit resource-close semantics. Without enforced close and drain constraints, resource reclamation becomes highly unpredictable — and in certain scenarios, leads to ClassLoaders that can never be collected.

Drawing on our exploration of runtime governance for long-lived JVM systems, we proposed a systematic ClassLoader lifecycle governance improvement plan to the Apache SeaTunnel community. The core shift is from passive “post-hoc residual reference hunting” to proactive “building deterministic reclaimable semantics and lifecycle closure.” We incrementally introduced explicit lifecycle close mechanisms, enforced classloading boundary constraints, and active residual reference cleanup. The governance proposal is currently under community discussion, with the Phase 1 optimization PR in review.

This talk will explore real-world community practices in Apache SeaTunnel, covering:

  1. Why long-lived systems cannot rely on implicit GC to manage underlying runtime resources
  2. How to build a ClassLoader governance standard for long-lived JVM systems and progressively land it in Apache open source projects
  3. How to complete a smooth, kernel-level architectural governance upgrade without breaking compatibility

Speaker

Jinxiang Yang | Creator of LingFrame & LingMirror | Apache SeaTunnel Community Member | JVM Runtime Governance

Jinxiang Yang is the creator of LingFrame (灵珑), an open source runtime governance framework for long-lived JVM single-process systems, and LingMirror, an IntelliJ IDEA plugin for static ClassLoader leak diagnosis. Focused on the governance of long-running JVM systems, he identified critical ClassLoader lifecycle gaps in Apache SeaTunnel through deep source code analysis, and proposed a systematic governance improvement plan that has been actively discussed in the SeaTunnel community. He believes a system is not complete when it runs — it is complete when its lifecycle can be observed, governed, and proven correct over time.

4. Session: Why SeaTunnel Engine Needs Clearer Distributed Abstractions

Time: August 9, 15:45–16:15 (China Standard Time)

Session Overview

In distributed data processing engines, responsibilities such as coordination, communication, and state management are often tightly coupled within a single framework. This can make the system difficult to evolve, extend, or replace individual components over time.

In this talk, I will share my experience working on the SeaTunnel engine, where we explored introducing clearer abstractions to reduce dependency on a single distributed framework. I will discuss why this kind of decoupling is necessary, the approach we considered, and the challenges we encountered along the way.

Rather than focusing on a specific technology choice, this talk will highlight the design questions and trade-offs involved in separating core responsibilities in a distributed engine.

Attendees will gain practical insights into how to think about abstraction boundaries and how to approach architectural evolution in real-world distributed systems.

Speaker

Doyeon Kim | Apache SeaTunnel Committer

Doyeon Kim is an Apache SeaTunnel Committer and a student with a strong interest in data engineering and distributed systems. She has contributed to SeaTunnel in areas such as connector development, engine improvements, and architectural discussions. Her recent work has focused on engine internals, dependency decoupling, and design trade-offs in distributed systems. Through this talk, she shares lessons learned from exploring clearer abstractions in the SeaTunnel engine.


💓 Don’t miss one of the biggest open source events of the year — Community Over Code Asia 2026. Join Apache SeaTunnel and connect with the global open source community!

We look forward to meeting you there!


🌟 Click Read More to register.

Limited seats are available. Register now and join us! 👆

Top comments (1)

Collapse
 
topstar_ai profile image
Luis Cruz

I'm particularly excited about the Data Lake & Data Warehouse track, where Lidong Dai will be sharing insights on AI-ready heterogeneous data integration and data governance, given his experience leading the development of WhaleStudio, a commercial data platform solution. The fact that Apache SeaTunnel will be covering key topics such as data integration, real-time synchronization, and AI-driven data pipelines resonates with my own experience working on similar projects, where I've seen the importance of leveraging open-source technologies to streamline data management and analytics. I'm curious to learn more about how the speakers will be addressing the challenges of implementing these technologies in production environments, and what best practices they'll be sharing from their own experiences. Will the sessions also be touching on the role of emerging technologies like serverless computing in the evolution of modern data infrastructure?