Every year, someone publishes a "State of Data Engineering" survey with 200 respondents, mostly from their own Slack community, and calls it representative. I ignore most of them. When Astronomer dropped their State of Apache Airflow 2026 report with 5,818 respondents across 122 countries, I paid attention. That's actual sample size.
I've spent years on both sides of the data engineering interview table, and I've watched candidates prep for tools that their target companies don't even use. The numbers in this survey confirm a few things I've seen firsthand, contradict some of the loudest takes on LinkedIn, and surface a gap that should make every data team uncomfortable.
What 5,800 Practitioners Are Actually Running in Production
The survey ran September 15 to November 20, 2025. 50 questions. 5,818 practitioners. 122 countries. Astronomer calls it the largest data engineering survey ever conducted. The 2025 edition had 5,250 respondents from 116 countries, so this is roughly 10% growth year over year. Steady accumulation.
One number jumped out immediately: Apache Airflow now has over 3,600 unique contributors. More than Spark's 2,000+. More than Kafka's 1,530. That's 135% more contributors than Kafka and 80% more than Spark. 3,600 humans who committed code to the project.
The contributor gap tells you something about the orchestration layer specifically. Airflow's ecosystem of provider packages (cloud connectors, custom operators, executor integrations) has lowered the friction for first-time contributors in a way that Spark's heavier core or Kafka's narrower domain focus hasn't matched. Contributor velocity is the best leading indicator of whether an open-source project will keep pace with the ecosystem around it. Right now, Airflow is outpacing both of its most obvious peers, and that matters more than any vendor's marketing slide.
84% Plan the Airflow 3 Migration. 26% Have Done It.
This gap is the most telling number in the entire report.
Airflow 3 shipped in April 2025. Less than a year later, 26% of respondents have completed the migration. 84% of those still on 2.x say they're planning it. That 58-point spread between planning and doing reflects engineering teams staring at a list of breaking changes and budgeting real project time.
The breaking changes aren't cosmetic. Airflow 3 removes SubDAGs entirely. It eliminates context variables like execution_date, prev_ds, and next_ds. Workers can no longer directly access the metadata database; everything routes through the REST API. xcom_pull(key=) stops searching upstream tasks. Each of these forces full DAG rewrites. If you've got 300 DAGs in production and half of them use execution_date for partition logic, you're looking at weeks of refactoring before you even touch your CI pipeline. That's why the gap exists.
Astronomer's own customer base tells a different story: 48% already run Airflow 3, and among their largest enterprise customers (50,000+ employees), 60% have deployed it. The gap between managed-platform customers and the broader community is a resource story. These companies have dedicated platform teams that can absorb migration work while the rest of the org keeps shipping. The 4-person data team at a Series B startup running 150 DAGs? They're going to plan the migration for Q3 and push it to Q4. Then Q1.
And the clock is ticking. Airflow 2 hit end of life in April 2026. No more security patches. No more bug fixes. No more provider package updates. If you're running Snowflake, Databricks, or BigQuery providers on Airflow 2, you're on borrowed time. Those vendors will start dropping 2.x compatibility in their own provider packages, and that second wave of pressure arrives 6 to 12 months after the official EOL. SOC 2, HIPAA, PCI-DSS; any certification that requires supported software makes this a compliance conversation on top of a technical one.
84% of teams know they need to migrate. 26% have done it. The gap is the actual cost of rewriting production DAGs while keeping the lights on.
If you're interviewing right now, expect questions on both Airflow 2 and 3 for at least another 12 months. The industry hasn't turned over yet.
The GenAI Production Gap
Here's where the survey gets uncomfortable.
Among all Airflow users, 32% report running GenAI or MLOps workloads in production. Among Astronomer's managed-platform customers, 62%. Among organizations that have been Astronomer customers for 2+ years, 83%.
51 percentage points between the general population and long-tenured platform customers. Read that again.
This has nothing to do with ambition. It has everything to do with infrastructure maturity compounding over time. Teams that spent years building out observability, error handling, schema governance, and pipeline idempotency can bolt on GenAI workloads because the plumbing already exists. Teams still figuring out how to make their batch jobs reliable aren't ready to orchestrate inference pipelines; they know it, and they're right to wait.
Separate research from K2view's 2026 enterprise survey backs this up: 76% of organizations cite data quality and consistency as their top barrier to production GenAI, and 62% say their enterprise data simply isn't ready. The bottleneck is the data platform underneath the model. MIT research referenced in the same analysis found 95% of enterprise AI pilots deliver zero P&L impact, which tells you the pilot-to-production gap is structural. You can't buy your way past it with a fancier model.
Meanwhile, 80%+ of the Astronomer survey respondents say they use AI tools to write pipelines but report that those tools hallucinate, lack context, and generate outdated syntax. I've been saying this for years: AI is making coding interviews increasingly pointless as a signal of engineering ability, but AI as a pipeline developer is nowhere close to replacing the engineer who knows why the pipeline broke at 3am last Tuesday and how to prevent it from happening again.
89% of Airflow users expect to expand orchestration into revenue-generating, external-facing solutions in 2026. The orchestration layer is moving from back-office utility to core product infrastructure. Your job as a data engineer just got a lot more visible to the business. That's both good and terrifying.
The Platform Split and What It Means for Your Next Data Engineering Interview
The conventional wisdom on LinkedIn says Databricks is running away with the market. The survey says: slow down.
Snowflake: 36.6%. Databricks: 34.7%. BigQuery: 27.8%. The total spread between first and third is under 9 percentage points. Snowflake leads Databricks by less than 2. Calling a winner here is irresponsible.
22% of Airflow users run 2 or more of the 3 major cloud data platforms. 1 in 5 teams operates a heterogeneous stack by choice. Intentional multi-platform architecture, not tech debt.
For interviews, this matters. A lot.
Snowflake loops emphasize the 3-layer architecture (storage, compute, cloud services), warehouse cost control, clustering strategies, and Snowpipe ingestion patterns. Databricks loops focus on Spark internals (shuffle semantics, partitioning), Delta Lake ACID transactions, and distributed systems thinking under concurrency pressure. These are fundamentally different prep surfaces. A single "data engineering interview" study plan that doesn't branch by platform will teach you the wrong material for half the roles you apply to.
Check the job posting. If it says Snowflake, drill Snowflake. If it says Databricks, drill Databricks. If it says both, they probably don't know what they want yet; ask in the recruiter screen. Don't waste 3 weeks studying the wrong platform because some influencer told you "Databricks is winning."
The concept-over-tool argument still holds, though. Data modeling, query optimization, understanding why things break; these transfer across all 3 platforms. The engineers who understand grain, normal forms, and late-arriving data will pass interviews on any platform after a week of syntax review. The ones who memorized COPY INTO but can't explain a slowly changing dimension will struggle regardless of which logo shows up in the job description. And if the loop includes a Python round, which most of them do now, python interview questions are what DataDriven is good for; we built the problem sets around the patterns that actually show up in production-focused loops, not LeetCode puzzles divorced from the job.
What This Survey Doesn't Tell You
Astronomer sponsors this survey and operates the platform it measures. The report blends community responses with "trends from Astro customer usage data," which means the 5,818 figure mixes voluntary respondents with inferred signals from their own product telemetry. Response rate and confidence intervals aren't published. The raw data isn't public. When the company running the survey also sells the managed version of the tool, numbers showing managed customers outperforming on every metric deserve some scrutiny.
Does that invalidate the findings? No. 5,800+ respondents across 122 countries is still the largest data engineering sample I've seen. The platform split, the migration gap, the GenAI adoption curve; these are directionally useful even with selection bias baked in. Use them as a compass, not a GPS coordinate.
The tools will keep changing every 18 months, like they always have. The problems won't. Schema drift, late-arriving data, upstream teams breaking contracts without telling you. Those are eternal. This survey confirms what I've watched play out for years: the industry is growing, the orchestration layer is becoming more strategic, and the gap between "I'm planning to do this" and "I've actually shipped it" is where careers get made.
What's your team actually running in production right now? Does it match what the survey says, or are you living in a completely different reality?
Top comments (0)