<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: DataDriven</title>
    <description>The latest articles on DEV Community by DataDriven (@datadriven).</description>
    <link>https://dev.to/datadriven</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3864671%2F923e8540-fa96-491d-adb6-0e01c42ec26a.png</url>
      <title>DEV Community: DataDriven</title>
      <link>https://dev.to/datadriven</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/datadriven"/>
    <language>en</language>
    <item>
      <title>Data Engineering Failures Cost Enterprises $3M a Month</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Thu, 01 Oct 2026 10:06:50 +0000</pubDate>
      <link>https://dev.to/datadriven/data-engineering-failures-cost-enterprises-3m-a-month-32h4</link>
      <guid>https://dev.to/datadriven/data-engineering-failures-cost-enterprises-3m-a-month-32h4</guid>
      <description>&lt;p&gt;Years ago I inherited a job that had been silently dropping around 40% of its records for 6 months. It didn't crash or alert, and the logs were clean. Someone in finance finally asked why the revenue numbers looked "a little light," and that's how we found out. The fix took an afternoon. Figuring out what had broken took 2 weeks, and the trust took a lot longer to rebuild.&lt;/p&gt;

&lt;p&gt;I bring this up because Fivetran just put a price tag on that kind of failure, and it's a big one. Their 2026 enterprise benchmark says pipeline failures expose large companies to roughly $3M a month. In the same stretch of 2026, companies cited AI as the reason for a growing share of tech layoffs. The work that prevents multimillion-dollar outages is the same work getting squeezed, and if you work in data engineering, that tension is also your best career lever right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Fivetran's 2026 benchmark found, and who it surveyed
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.fivetran.com/blog/the-enterprise-data-infrastructure-benchmark-report-2026" rel="noopener noreferrer"&gt;Fivetran's Enterprise Data Infrastructure Benchmark Report 2026&lt;/a&gt; surveyed 500 senior data and technology leaders in Q4 2025. Every respondent works at an organization with 5,000+ employees, across the US, UK, EMEA, and APAC. The industry mix leans toward financial services (23%), manufacturing (20%), and tech (20%), followed by retail/CPG, healthcare, and a little hospitality. Fivetran reports 95% confidence with a ±4.4% margin of error.&lt;/p&gt;

&lt;p&gt;Here are the headline numbers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;4.7 pipeline breaks per month on average, rising to 8.3 at the largest enterprises&lt;/li&gt;
&lt;li&gt;60.4 hours of downtime per month, with each incident taking nearly 13 hours to resolve&lt;/li&gt;
&lt;li&gt;53% of engineering capacity spent on maintenance instead of new work&lt;/li&gt;
&lt;li&gt;$2.2M per year per enterprise in maintenance labor alone&lt;/li&gt;
&lt;li&gt;An average shop runs 328 pipelines with 35 full-time engineers&lt;/li&gt;
&lt;li&gt;97% of leaders said pipeline failures have slowed their analytics or AI programs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The math holds together, which I appreciate. 4.7 breaks at roughly 13 hours each comes to roughly 61 hours, which matches the 60.4. Fivetran's &lt;a href="https://www.fivetran.com/press/data-pipeline-failures-cost-enterprises-3-million-per-month-fivetran-benchmark-finds" rel="noopener noreferrer"&gt;press release&lt;/a&gt; puts business exposure at $49,600 per hour of downtime. Multiply that by 60.4 hours and you land right around $3M a month. A single incident can reach $1.4M.&lt;/p&gt;

&lt;p&gt;Now the part a vendor's marketing team won't put in bold: Fivetran sells managed ELT. The report says legacy and DIY systems break 30% to 47% more often than managed ones, and that managed-platform adopters are "nearly 2x as likely" to beat ROI expectations. That might be true. It's also exactly the finding you'd expect from a company whose business is managed pipelines. Treat the breakage and downtime counts as decent signal and the "buy managed and your problems go away" conclusion as a sales pitch.&lt;/p&gt;

&lt;p&gt;I've migrated legacy warehouses at 3am. Managed connectors do remove a whole category of pain, mostly API pagination, auth token rotation, and schema changes in SaaS sources. They don't fix a bad data model, an upstream team renaming a column without telling anyone, or a join that quietly fans out and doubles your revenue.&lt;/p&gt;

&lt;p&gt;The number I keep coming back to is the 53%. Of 35 engineers, around 18 spend their week keeping existing pipelines alive. Every senior DE I know would call that number low.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pipeline failures and layoffs in the same month
&lt;/h2&gt;

&lt;p&gt;Now set that next to the layoff data.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.challengergray.com/wp-content/uploads/2026/06/Challenger-Report-May-2026.pdf" rel="noopener noreferrer"&gt;Challenger, Gray &amp;amp; Christmas May 2026 report&lt;/a&gt; has AI cited in about 22% of all US job cuts through May. In May alone it was 40%. Later 2026 coverage of Challenger's numbers puts the year-to-date share around 24% through July, with AI as the most-cited reason for several months in a row. The cumulative counts in secondary write-ups don't fully agree with each other, so I won't pretend to know the exact figure, but the direction is clear. Anyone still quoting "roughly a fifth" is out of date, because the share rose through mid-year.&lt;/p&gt;

&lt;p&gt;TechCrunch has kept a running list of large tech layoffs where employers cited AI, including Amazon, Oracle, Microsoft, Meta, and PayPal. Andy Challenger summed it up this way: "Tech remains the center of gravity for this year's cuts, and AI is still the reason companies give."&lt;/p&gt;

&lt;p&gt;Keep the phrase "the reason companies give" in mind. What a company says in a press release about a layoff and what actually caused it are often unrelated. I've survived multiple layoff waves. The official reason has always been a story told to the board, and this year the story is AI.&lt;/p&gt;

&lt;p&gt;The economics don't work, and I mean that literally. A typical enterprise in Fivetran's sample has $3M a month of downtime exposure, and the people preventing it are a line item in the maintenance budget. Cut 5 of those engineers to save about $1M a year in salary, and you only need a few extra 13-hour incidents to give all of it back. You're also teaching the business that the dashboards can't be trusted, which costs more than any of it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The labor that keeps $3M a month of downtime contained is the easiest line item to cut and the most expensive one to lose.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I'll be fair to the other side. The same research also shows data engineering postings growing, with demand moving toward people who can handle infrastructure, governance, and cost control. The squeeze lands hardest on generic and lower-tier roles. So the DE role isn't disappearing. It's being repriced, and the people doing the repricing don't fully understand what the job involves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the maintenance hours go
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://montecarlo.ai/blog-data-quality-survey" rel="noopener noreferrer"&gt;Monte Carlo's 2026 State of Data Quality survey&lt;/a&gt; answers most of it. 68% of teams need 4+ hours just to detect an incident, up from 62% in 2022. Average resolution time is 15 hours. And the stat that should embarrass everyone in this industry: 74% of data issues are found by business stakeholders first, up from 47% in 2022.&lt;/p&gt;

&lt;p&gt;So for 3 out of 4 issues, the monitoring system is a VP in a Slack channel asking why the numbers look weird. That was my 40% record-drop job exactly, except today it happens at 3x the scale.&lt;/p&gt;

&lt;p&gt;That's where the 53% goes. Most of it is detective work: finding which of the 328 pipelines broke, which upstream source changed, which backfill needs to rerun, and which downstream tables are now wrong. The actual fix is usually the shortest part of the incident.&lt;/p&gt;

&lt;p&gt;The other surveys in the research fill in the rest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;dbt's 2026 State of Analytics Engineering report found 72% of data teams prioritize AI coding, while only 24% prioritize AI-assisted pipeline testing and observability. So we're speeding up how fast people write pipelines and barely investing in knowing when they break.&lt;/li&gt;
&lt;li&gt;Atlassian's DX report for Q2 2026 found AI saves engineers 4 to 6 hours a week, yet the innovation ratio barely moved, staying roughly 57% to 58%. The freed-up time gets absorbed by the maintenance backlog.&lt;/li&gt;
&lt;li&gt;Teams with automated observability reportedly resolve incidents about 4x faster. That claim is secondhand, so hold it a bit loosely, but every incident I've worked agrees with the direction.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tooling is catching up, for what it's worth. Apache Airflow 3.3 shipped AIP-103, a first-class task and asset state store. It persists key-value state across worker crashes and retries, which replaces the XCom hacks that never survived a retry. Apache Iceberg 1.12 replaces position delete files with deletion vectors, compact bitmaps that Dremio says cut read latency by 50% to 80% for high-frequency deletes. That matters a lot for CDC and GDPR workflows.&lt;/p&gt;

&lt;p&gt;Both are good releases. Neither changes the fundamentals. Iceberg still needs compaction, snapshot expiration, orphan cleanup, and manifest optimization running in a coordinated way, and if you get orphan retention wrong a mid-write failure can silently corrupt your table. A state store won't help if you don't understand idempotency, because you'll just persist the wrong state faster.&lt;/p&gt;

&lt;p&gt;This is the concepts-over-tools argument again. Idempotency, late-arriving data, schema contracts, grain, and backfill strategy work the same in Airflow 2, Airflow 3.3, Dagster, or a cron job someone wrote in 2014. Tools change every 18 months, and the reasons pipelines fail don't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your next data engineering interview
&lt;/h2&gt;

&lt;p&gt;Compensation hasn't caught up with reliability work yet. &lt;a href="https://hiringlab.indeed.com/2026/09/17/ai-exposure-isnt-squeezing-advertised-pay-in-the-us-its-boosting-it/" rel="noopener noreferrer"&gt;Indeed Hiring Lab's September 2026 analysis&lt;/a&gt; found advertised pay in highly AI-exposed occupations up roughly 46% since 2021, compared with 25% for less-exposed fields. The AI premium on identical job titles is about 4.7%, and much larger at senior levels. Mid-level DE base pay sits about $139K median. Meanwhile the people preventing $3M a month in downtime get paid like plumbers. (Plumbers, to be fair, are often doing fine.)&lt;/p&gt;

&lt;p&gt;I don't think that lasts, and the hiring signals already show it shifting. Recruiting guides for 2026 now treat "reliability signals" as the minimum: SLAs, monitoring, alerting, backfills, retries, tests, incident response, and postmortems. Senior job postings name observability architecture as a core responsibility, and some staff-level postings run past $260K.&lt;/p&gt;

&lt;p&gt;Here's how to use that.&lt;/p&gt;

&lt;p&gt;Rewrite your resume around incidents. "Built ETL pipelines using Airflow and Snowflake" is a tool list. "Cut detection time on a revenue pipeline from next-day to 20 minutes by adding row-count and freshness checks" is a story. Don't tell me you "ensured data reliability across a wide array of mission-critical workflows." If I read one more bullet like that I'm putting my fist through drywall.&lt;/p&gt;

&lt;p&gt;Have a postmortem ready. Interviewers increasingly ask you to walk through a real incident: how you detected it, what the blast radius was, how you fixed it, and what you changed so it couldn't happen again. Pick your best one and rehearse it until you can tell it in 3 minutes without rambling. If you don't have one, you haven't been doing this long enough, or you were lucky. Either way, go break something in a side project and fix it.&lt;/p&gt;

&lt;p&gt;Prep for pipeline architecture. I've watched people with 10 YOE get downleveled because they couldn't explain why their design would survive a late-arriving partition. DEs don't need to whiteboard load balancers. You need to explain how your pipeline handles retries without double-writing, how you'd backfill 90 days without taking down the warehouse, and where your quality checks go. That's the gap we built DataDriven to cover, so for pipeline design practice, try DataDriven: it drills exactly those failure-mode questions, which is where loops separate seniors from everyone else.&lt;/p&gt;

&lt;p&gt;Ask about the 53% in your interviews. Ask every hiring manager how many incidents they had last quarter and who found them first. If the answer is "the business usually tells us," you're interviewing for a firefighting job, so price the offer accordingly. If they actually know their numbers, that's a team worth joining.&lt;/p&gt;

&lt;p&gt;If you're worried about layoffs, become the person who knows why things break. Every layoff list I've seen spared the engineer who held the context on the scary pipelines. It comes down to what happens when the CFO asks, "who can tell me why the numbers are wrong?" and only one name comes up.&lt;/p&gt;

&lt;p&gt;I've been through 3 waves of "data engineering is getting automated away." I'm still here and still debugging the same categories of problems. The Fivetran numbers just put a dollar figure on what every on-call DE already knew: the hard part of the job is keeping things working. The companies cutting that work now will rehire for it at a premium after their first $1.4M incident, and the engineers who can tell a good postmortem story will get those offers.&lt;/p&gt;

&lt;p&gt;So here's my question: in your last pipeline incident, who noticed first, your monitoring or someone from the business?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>career</category>
      <category>devops</category>
      <category>database</category>
    </item>
    <item>
      <title>The Real Data Behind the Data Engineering Hiring Panic</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 29 Sep 2026 10:08:43 +0000</pubDate>
      <link>https://dev.to/datadriven/the-real-data-behind-the-data-engineering-hiring-panic-3dhh</link>
      <guid>https://dev.to/datadriven/the-real-data-behind-the-data-engineering-hiring-panic-3dhh</guid>
      <description>&lt;p&gt;Last month 3 different people sent me the same screenshot. It was a slick graphic claiming tech lost 150K jobs this year while data engineering grew 414%, and entry-level postings "collapsed 67%." Each of them asked me whether they should panic, pivot, or quit.&lt;/p&gt;

&lt;p&gt;I went looking for where those numbers come from. I couldn't find a dataset behind any of them.&lt;/p&gt;

&lt;p&gt;Meanwhile, 3 real sources published numbers this year that you can check yourself: Stanford's payroll research, LinkedIn's own ranking of fast-growing jobs, and the Bureau of Labor Statistics. What they show is narrower and more specific than the viral numbers, and for anyone planning a data engineering career in 2026 it's the more useful picture. Early-career hiring got squeezed hard. Experienced people are doing fine. The field itself is growing slowly and steadily.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the viral data engineering hiring numbers come from
&lt;/h2&gt;

&lt;p&gt;I've been in this industry long enough to see 3 waves of "data engineering is getting automated away." Each wave came with its own scary chart, and each chart had the same flaw: nobody could tell you where the numbers came from.&lt;/p&gt;

&lt;p&gt;The current batch fits that pattern. The posts pushing "414% growth" and "67% collapse" (some of them right here on DEV) point to "industry reports" and "Gartner" in general terms. They don't link anything, name a dataset, or give a sample size. The closest thing I found to a source for the 67% figure is a job-postings analytics vendor's blog that says entry-level DE postings fell 67% between October 2023 and November 2024. It cites proprietary scraped posting data that you can't audit.&lt;/p&gt;

&lt;p&gt;That doesn't prove the number is wrong. What it does mean is that nobody can check it, and when a number can't be checked, you shouldn't be making career decisions off it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A statistic with no dataset behind it is a vibe with a percent sign.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What bugs me most is that these numbers travel. They get screenshotted, stripped of any context, and dropped into career advice threads, where someone with 2 years of experience reads them at 11pm and decides to drop out of their job search. I've watched it happen, and I've been that person at 11pm myself. I did 20-some loops in one job search. I didn't need a fake statistic telling me the market was over when the real market was already beating me up just fine.&lt;/p&gt;

&lt;p&gt;So let's look at the numbers you can actually verify.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stanford's payroll data on early-career hiring
&lt;/h2&gt;

&lt;p&gt;The Stanford Digital Economy Lab worked with ADP to track real payroll records: millions of actual paychecks across tens of thousands of firms, instead of job postings or surveys. The research is called &lt;a href="https://digitaleconomy.stanford.edu/publication/canaries-in-the-coal-mine-six-facts-about-the-recent-employment-effects-of-artificial-intelligence/" rel="noopener noreferrer"&gt;"Canaries in the Coal Mine"&lt;/a&gt;, and its headline finding is that employment of software developers aged 22 to 25 fell about 20% from its late 2022 peak.&lt;/p&gt;

&lt;p&gt;Over the same period, developers aged 30 and older at the same kinds of firms, in the same roles, grew employment by 6 to 12%. So it's an age effect inside a single occupation. Juniors lost ground while experienced people gained it, and total demand for developers didn't fall off a cliff.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://digitaleconomy.stanford.edu/news/canariesaug26/" rel="noopener noreferrer"&gt;August 2026 update&lt;/a&gt; extends the finding past developers. Workers aged 22 to 25 across all highly AI-exposed occupations now sit 19% below where they'd be if they'd kept pace with less-exposed peers. In July 2025 that gap was 15%, so it's getting wider.&lt;/p&gt;

&lt;p&gt;The detail I think matters most is how the decline happens. Stanford found that the adjustment runs "primarily through reduced hiring of young workers rather than increased separations." Separation rates actually fell, including for young workers in exposed fields. Firms aren't firing juniors in large numbers. They've just stopped opening the door for new ones.&lt;/p&gt;

&lt;p&gt;That fits what I'm hearing from hiring managers. None of them describe a "fire the juniors" program. What happens instead is that a req for a junior slot quietly gets converted to a senior one, or backfills get put off until the team decides it can live without the role.&lt;/p&gt;

&lt;h3&gt;
  
  
  The caveats Stanford's own critics raise
&lt;/h3&gt;

&lt;p&gt;This research has weak spots and you should know about them.&lt;/p&gt;

&lt;p&gt;The Fed started its most aggressive rate-hiking cycle in 40 years in March 2022. Job postings in AI-exposed sectors started dropping around then, about 8 months before ChatGPT launched. Some economists, including a group writing in &lt;a href="https://www.promarket.org/2026/06/17/ai-is-not-reducing-employment-but-rather-who-gets-hired/" rel="noopener noreferrer"&gt;ProMarket&lt;/a&gt;, argue the data fits an interest-rate hiring freeze about as well as it fits an AI story. Stanford's side points out that an occupation's AI exposure doesn't correlate with how sensitive it is to interest rates, which suggests these are 2 separate channels.&lt;/p&gt;

&lt;p&gt;My read is that it's probably both, and nobody can cleanly separate them yet. You'll also see some secondhand write-ups quoting different figures for the developer decline, so go to Stanford's own dashboard instead of trusting the reposts.&lt;/p&gt;

&lt;p&gt;The other caveat matters a lot for our field. Stanford measures "software developers," and "data engineers" aren't a separate category in it. Applying this finding to DE specifically is an inference. I think it's a reasonable one, because junior DE work looks a lot like junior SWE work, but it's still an extrapolation.&lt;/p&gt;

&lt;h2&gt;
  
  
  LinkedIn and BLS on data engineering growth
&lt;/h2&gt;

&lt;p&gt;Next is the "414% growth" claim.&lt;/p&gt;

&lt;p&gt;LinkedIn publishes a yearly list called Jobs on the Rise, which ranks the fastest-growing roles using its own members' job transitions. The &lt;a href="https://www.linkedin.com/pulse/linkedin-jobs-rise-2026-25-fastest-growing-roles-us-linkedin-news-dlb1c" rel="noopener noreferrer"&gt;2026 edition&lt;/a&gt; looked at jobs started between January 2023 and July 2025. AI Engineer is #1, AI Consultant/Strategist is #2, New Home Sales Specialist is #3 (really), and Data Annotator is #4.&lt;/p&gt;

&lt;p&gt;Data engineer doesn't show up anywhere in the top 25. It wasn't in 2025's top 25 or 2024's either.&lt;/p&gt;

&lt;p&gt;If DE had grown 414%, LinkedIn would have noticed, because tracking this is the whole point of the list. So a role supposedly growing 4x shows up 0 times in 3 years of the one ranking built to catch fast growth. I'll let you decide which of those numbers to believe.&lt;/p&gt;

&lt;p&gt;Missing from a &lt;em&gt;fastest-growing&lt;/em&gt; list doesn't mean shrinking, though. It means DE is a mature, established role growing at a normal pace, and the explosive growth went to titles that barely existed 3 years ago. I'd honestly be more nervous if DE were on the list, since hype-driven titles tend to crater just as quickly.&lt;/p&gt;

&lt;h3&gt;
  
  
  The BLS baseline
&lt;/h3&gt;

&lt;p&gt;The federal government doesn't track "data engineer" at all. The Standard Occupational Classification system groups workers by the work they do, not by job title, so data engineers get spread across database architects, software developers, and data scientists depending on what their job looks like day to day. After more than a decade of the title being everywhere in industry, it still doesn't have its own code.&lt;/p&gt;

&lt;p&gt;The closest proxy is Database Administrators and Architects, and the &lt;a href="https://www.bls.gov/ooh/computer-and-information-technology/database-administrators.htm" rel="noopener noreferrer"&gt;BLS projects 4% growth&lt;/a&gt; for it over the next decade, "about as fast as the average for all occupations." That's roughly 7,300 openings a year, mostly from people retiring or leaving.&lt;/p&gt;

&lt;p&gt;The 4% average hides a split inside the category. Traditional database administrators are projected at 0% growth. Database architects are projected at 9%. Software developers, where plenty of DEs get counted, are projected at 10%. Data scientists are at 35%.&lt;/p&gt;

&lt;p&gt;Data engineering lands somewhere between 4% and 10%, depending on how you count, and no official number pins it down more tightly than that. That growth is boring and positive, which is how a healthy, established profession looks.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;414% was never real. 4 to 10% is. Boring growth is still growth, and it compounds for people who stay in the game.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What this means for your data engineering interview strategy
&lt;/h2&gt;

&lt;p&gt;Put the 3 sources together and here's what I think they show, with the caveat that this is my interpretation and no single dataset proves it. The door is narrower at the bottom: fewer junior seats are opening, and whoever does get hired has to show more on day one. The middle and top are holding up, with experienced people in the same roles gaining employment. And teams aren't being gutted. Separations are low, so the pressure falls on who gets hired next, and current employees aren't being pushed out.&lt;/p&gt;

&lt;p&gt;Your strategy depends on which side of that line you're on.&lt;/p&gt;

&lt;h3&gt;
  
  
  If you're early-career
&lt;/h3&gt;

&lt;p&gt;This is the hard group, and I won't pretend it isn't. A 20% drop in early-career developer employment is a real hit, and the numbers back it up. If you've been getting rejected over and over, a big part of that is the market. It says little about how capable you are.&lt;/p&gt;

&lt;p&gt;Still, "fewer junior seats" means you have to look like you've already done the job. Companies are cutting the roles where someone gets paid to learn, so show that you've learned already. Build something that breaks. Run a pipeline on a schedule for 3 months and write down every time it failed and what you did about it. Skip the tutorial DAG. The actual job is mostly debugging, things like working out why 2M rows silently disappeared last Tuesday, and that's the story that gets a junior noticed.&lt;/p&gt;

&lt;p&gt;Also, drop the pure-SWE framing. Data modeling is where junior DE candidates get separated from everyone else. An AI tool can write a medium LeetCode solution in a few seconds, which makes that a weak signal, and interviewers increasingly know it. What still tells an interviewer something is whether you can take a messy business process and design a model for it: define the grain, handle late-arriving data, and explain why you picked one approach over another. The prep I keep sending people to is the same one I'd use myself: for data modeling interview questions, try datadriven, since its problems are built around real pipelines and schemas instead of puzzle trivia.&lt;/p&gt;

&lt;h3&gt;
  
  
  If you've got a few years in
&lt;/h3&gt;

&lt;p&gt;You're on the side of the data that's growing, so spend less time on doom threads and more on pricing yourself properly.&lt;/p&gt;

&lt;p&gt;The interview is where experienced people lose money. I've watched engineers with 10 years of experience get downleveled because they couldn't explain their own design decisions when put on the spot. They did the work; they just couldn't tell the story of it in 45 minutes. That's a skill you can practice, completely separate from being good at the job.&lt;/p&gt;

&lt;p&gt;Get your war stories straight. For every big project, know the grain of the model, what broke, what it cost the business, and what you'd change now. Put the vague resume language in the trash. "Leveraged modern data stack to drive stakeholder value" tells me nothing. "Migrated 400 tables with zero downtime and cut the nightly batch from 6 hours to 90 minutes" tells me you're senior.&lt;/p&gt;

&lt;h3&gt;
  
  
  If you're hiring
&lt;/h3&gt;

&lt;p&gt;This is aimed at the people on my side of the table. The Stanford researchers raise the uncomfortable question here: if nobody hires juniors today, who are the seniors in 2032? Every senior DE I know started as someone who broke production at least once while someone else paid their salary. Cutting junior hiring saves money this quarter, and the cost shows up years later as a senior shortage you can't hire your way out of.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the job market like a data engineer
&lt;/h2&gt;

&lt;p&gt;We're supposed to be the people who check the data. We trace a number back to its source table before we put it on a dashboard. We don't trust a metric until we know its grain, its filters, and who calculated it.&lt;/p&gt;

&lt;p&gt;Do the same with job market content. When you see a scary number, ask the questions you'd ask about a broken pipeline: what's the source, what's the population, what's the time window, and is it measuring postings, hires, or actual paychecks? Postings are noisy, hires are better, and payroll records are the closest thing we have to ground truth.&lt;/p&gt;

&lt;p&gt;When I ran that check on this year's data, a narrower story came out than the viral posts tell. Early-career hiring really is tight. Experienced engineers are gaining ground. The field is growing at a steady, unremarkable pace that no government agency can measure precisely, because government agencies still don't know what we're called.&lt;/p&gt;

&lt;p&gt;I've been through enough hype cycles to recognize this one. In a few years someone will need to debug why the new AI pipeline is silently dropping records, and I'd put money on it being a data engineer.&lt;/p&gt;

&lt;p&gt;So here's my question: when did you last change a career decision because of a job-market statistic, and did you ever check where that number came from?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>career</category>
      <category>interview</category>
      <category>ai</category>
    </item>
    <item>
      <title>dbt's 2026 Survey: Data Engineers Chase AI, Skip the Testing</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Thu, 24 Sep 2026 10:08:01 +0000</pubDate>
      <link>https://dev.to/datadriven/dbts-2026-survey-data-engineers-chase-ai-skip-the-testing-d96</link>
      <guid>https://dev.to/datadriven/dbts-2026-survey-data-engineers-chase-ai-skip-the-testing-d96</guid>
      <description>&lt;p&gt;The worst bug I ever inherited was a job that silently dropped 40% of records for 6 months before anyone noticed. No test caught it and no alert fired. Nobody owned the target table, so nobody looked. The code ran green every night and did the wrong thing every night.&lt;/p&gt;

&lt;p&gt;I think about that job a lot lately, because the industry has decided to build a lot more of them and build them faster.&lt;/p&gt;

&lt;p&gt;dbt Labs' &lt;a href="https://www.getdbt.com/resources/state-of-analytics-engineering-2026" rel="noopener noreferrer"&gt;2026 State of Analytics Engineering report&lt;/a&gt; surveyed 363 practitioners and leaders between December 2025 and February 2026. 72% of teams now prioritize AI-assisted coding. Only 24% prioritize AI-assisted pipeline management, meaning testing and observability. So 3 times as many teams are working on shipping code faster as are working on checking it.&lt;/p&gt;

&lt;p&gt;The code is the cheap part now, and checking it is where the value is. Teams that spend on generation and skip the guardrails are running up debt that someone like me will get paid to clean up in 2028.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers from dbt's 2026 survey
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Trust in data as a priority went from 66% to 83% in a single year. According to &lt;a href="https://www.getdbt.com/blog/new-dbt-labs-report-finds-ai-driven-acceleration-is-outpacing-trust-and-governance" rel="noopener noreferrer"&gt;dbt Labs' own writeup&lt;/a&gt;, that's the steepest single-year increase in the survey's history.&lt;/li&gt;
&lt;li&gt;Speed as a priority rose from 50% to 71%.&lt;/li&gt;
&lt;li&gt;Cost reduction barely moved, from 48% to 53%.&lt;/li&gt;
&lt;li&gt;71% worry about hallucinated or incorrect outputs reaching stakeholders.&lt;/li&gt;
&lt;li&gt;41% say ambiguous data ownership is still a problem, roughly flat from last year.&lt;/li&gt;
&lt;li&gt;57% report higher warehouse and compute spend, while only 36% report bigger team budgets.&lt;/li&gt;
&lt;li&gt;77% of managers emphasize AI for productivity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Put those in a row and you get teams that want to go faster, say they care much more about correctness, and are putting their effort into the speed half.&lt;/p&gt;

&lt;p&gt;To be fair to the respondents, wanting speed and trust together is reasonable. Every stakeholder I've ever had wanted both. What bugs me is the ratio. Speed moved 21 points, trust moved 17, and the tooling investment went overwhelmingly to speed.&lt;/p&gt;

&lt;p&gt;The report tells you trust jumped, but it doesn't say why, and it names no single incident or regulation behind it. My guess, and it's only a guess, is that enough teams shipped an AI-generated model that quietly produced a wrong number in front of an executive. Once that happens, trust becomes very important very fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  Data trust jumped 17 points and nobody bought tests
&lt;/h2&gt;

&lt;p&gt;Trust is a priority for 83% of respondents, but only 24% put resources into the machinery that produces it.&lt;/p&gt;

&lt;p&gt;I've been in those planning meetings, and the gap comes down to economics.&lt;/p&gt;

&lt;p&gt;AI-assisted coding pays off right away and is easy to see. An engineer turns a 2-day model into a 2-hour model, their manager sees a burndown chart drop, and everyone feels productive. Data observability pays off in incidents that never happen. Nobody holds a retro for the dashboard that stayed correct. You can't put "nothing broke" on a slide and get more budget.&lt;/p&gt;

&lt;p&gt;The compute numbers make it worse. 57% of teams are spending more on the warehouse, while only 36% got more money for people. Compute bills per query and scales on demand. Headcount needs approval, a req, a hiring loop, and 3 months of ramp. When you're squeezed, you let the warehouse absorb the inefficiency and let the backlog absorb the testing.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;AI made writing a dbt model nearly free. Nobody made verifying one free. Whatever you ship without a test turns into a liability with a slow fuse.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I've been through 3 waves of "data engineering is getting automated away." Each time the tools changed and the failure categories didn't. You still get schema drift, late-arriving data, upstream teams breaking contracts without telling you, and joins that fan out and silently double your revenue. AI gets you to those failures faster. It doesn't prevent them.&lt;/p&gt;

&lt;p&gt;What worries me most about AI-generated SQL is what it does to review. When a junior engineer writes a bad join, it usually looks bad. The CTE names are weird, the formatting is off, and something makes a reviewer squint. AI-generated SQL looks clean and confident whether it's right or wrong. Review has always relied partly on the author's code looking unsure, and that signal is gone.&lt;/p&gt;

&lt;p&gt;If you're on a team in the 72% but not the 24%, you can do most of this yourself without a platform initiative.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Test at the grain. Every model gets a uniqueness test on its declared grain. If you can't state the grain in a sentence, the model isn't done.&lt;/li&gt;
&lt;li&gt;Test row counts across boundaries. Rows in versus rows out on every join that shouldn't change cardinality. That's the test that would have caught my 40% drop in week 1 instead of month 6.&lt;/li&gt;
&lt;li&gt;Put freshness checks on anything a human looks at. Stale data that looks fresh does more damage than an obviously broken dashboard.&lt;/li&gt;
&lt;li&gt;Make AI-generated models carry tests before merge. If the model wrote the SQL, it can write the tests too. Review the tests harder than the SQL.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is new, and dbt has made most of it trivial for years. The survey says most teams still aren't prioritizing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ownership problem data engineering keeps dodging
&lt;/h2&gt;

&lt;p&gt;The number I keep coming back to is the 41% with ambiguous ownership, because it didn't move. According to dbt Labs, technical integration barriers dropped from 35% to 27% year over year, so the tooling got better. Ownership stayed stuck.&lt;/p&gt;

&lt;p&gt;Other surveys point the same way. Joe Reis' &lt;a href="https://joereis.substack.com/p/the-2026-state-of-data-engineering" rel="noopener noreferrer"&gt;2026 State of Data Engineering survey&lt;/a&gt; of more than 1,000 data engineers found 59% citing pressure to move fast and 51% citing lack of ownership as top pain points. Only 11% said their data modeling was going well.&lt;/p&gt;

&lt;p&gt;Then there's the result I'd put on a billboard. The Practical Data Community's &lt;a href="https://practicaldatamodeling.substack.com/p/april-2026-pdc-state-of-data-modeling" rel="noopener noreferrer"&gt;April 2026 data modeling survey&lt;/a&gt; asked 334 people what would most improve their modeling. Training came first at 28.1%, then clearer business requirements at 24.6%, more time at 21.6%, and dedicated ownership at 21%. Better tooling got 4.8%.&lt;/p&gt;

&lt;p&gt;Vendors keep selling tools to an industry that has told them, in writing, that tools aren't the bottleneck.&lt;/p&gt;

&lt;p&gt;The same survey found 42.5% of data models are owned by whoever built the pipeline, 19.2% have a dedicated modeler or architect, and 7.8% have no formal owner at all. I've worked in all 3 setups. The "whoever built it owns it" model works until that person leaves, gets reorged, or gets laid off. After that the table becomes folklore.&lt;/p&gt;

&lt;p&gt;Now add AI to that. When an engineer writes a model by hand, they at least know what they meant. When an agent generates it from a prompt, the "owner" is whoever typed the prompt, and they may not be able to explain the join logic in the output. The result is an untested model that no human wrote and nobody is named as owning, feeding an executive dashboard. It's the setup where hallucinated numbers reach stakeholders, and 71% of respondents already worry about exactly that.&lt;/p&gt;

&lt;p&gt;Pooja Crahen, a senior manager of analytics engineering at Okta, said in the dbt Labs release that you can't get both speed and trust without discipline in modeling, validation, and ownership, and that the discipline has to be a requirement rather than a best practice. That's her opinion, and I agree with it. I'd add that it has to be a requirement someone is accountable for, with a name next to it.&lt;/p&gt;

&lt;p&gt;AI governance frameworks won't fix this on their own. I've seen plenty of governance programs that were thorough on paper and did nothing in practice. What fixes it is a person who can say "no, this doesn't ship without a grain test" and be listened to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Using the trust gap in your next data engineer interview
&lt;/h2&gt;

&lt;p&gt;The gap is a career opportunity if you position yourself right.&lt;/p&gt;

&lt;p&gt;I've been on both sides of the hiring table. Everyone now has AI coding assistants, so writing SQL quickly no longer sets you apart. I've done 20+ loops in a single job search, and I've sat on panels where we passed on strong candidates for dumb reasons. The candidates who stood out most, even before this survey, were the ones who could explain how they knew their pipeline was right.&lt;/p&gt;

&lt;p&gt;Senior interviews are already testing for this through behavioral questions. "Tell me about a time you caught a data issue before anyone else noticed." "Tell me about a pipeline you improved without being asked." Those are ownership questions. They check whether you treat data quality as your job or as something downstream.&lt;/p&gt;

&lt;p&gt;Most candidates answer them badly. They talk about the pipeline they built. The strong answer covers what broke, how they found it, what the blast radius was, and what they put in place so it can't happen again. The actual job is debugging, and the interviewers who've done the job know it.&lt;/p&gt;

&lt;p&gt;So here's the prep plan.&lt;/p&gt;

&lt;p&gt;Get a debugging war story and tell it with numbers. Say "Found a fan-out join inflating revenue 12% for 3 weeks; added grain tests to 40 models," not "improved data quality."&lt;/p&gt;

&lt;p&gt;Own the modeling round. Know grain, slowly changing dimensions, and fact table design well enough to explain why a model breaks as well as how to build it. Data modeling is the skill AI is worst at faking, because it depends on business context the model doesn't have. We're biased because it's ours, but for snowflake interview questions, try datadriven.io. We built it around grain, dedup, and "why is this number wrong" style problems.&lt;/p&gt;

&lt;p&gt;Talk about observability like someone who's been paged. Know freshness, volume and schema checks, and know which alerts get looked at and which ones get muted by week 2. Interviewers can tell whether you've been on call or only read about it.&lt;/p&gt;

&lt;p&gt;Talk about AI without being breathless or dismissive. "I use it to draft models and I make it write tests I review harder than the SQL" is a senior answer. "AI writes all my code" and "I don't trust AI" are both junior answers.&lt;/p&gt;

&lt;p&gt;This is where the economics favor you. Teams are spending more on compute than on people, generating code faster than they can review it, and 83% of them say trust is suddenly a top priority. The engineer who can turn "we care about trust" into tests, owners, and alerts is solving the exact problem those teams have admitted to in writing.&lt;/p&gt;

&lt;p&gt;So I'm curious where your team lands: are you in the 24% actually investing in testing and data observability, or are you shipping AI-generated models and hoping the dashboards stay honest?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>ai</category>
      <category>career</category>
      <category>database</category>
    </item>
    <item>
      <title>Spark 4.x Is Stable. The Data Engineering Upgrade Tax.</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 22 Sep 2026 10:10:55 +0000</pubDate>
      <link>https://dev.to/datadriven/spark-4x-is-stable-the-data-engineering-upgrade-tax-44io</link>
      <guid>https://dev.to/datadriven/spark-4x-is-stable-the-data-engineering-upgrade-tax-44io</guid>
      <description>&lt;p&gt;Spark 4.2.0 shipped July 14, 2026. If you're still running 3.x in production, the grace period is over.&lt;/p&gt;

&lt;p&gt;I've been through enough platform migrations to know the pattern. The release drops. Everyone says "we'll get to it next quarter." Then 18 months later you're running an unsupported runtime with 3 engineers who know how to keep it alive and a Confluence page titled "DO NOT TOUCH" as your only documentation. The &lt;strong&gt;Apache Spark&lt;/strong&gt; 3.x to 4.x jump is that migration, happening right now, and the breaking changes are worse than most teams realize.&lt;/p&gt;

&lt;p&gt;Here's what the &lt;strong&gt;spark upgrade&lt;/strong&gt; actually costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  ANSI Mode: The Silent Pipeline Killer
&lt;/h2&gt;

&lt;p&gt;The biggest breaking change in Spark 4.0 is a default that flipped.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;spark.sql.ansi.enabled&lt;/code&gt; is now &lt;code&gt;true&lt;/code&gt;. In Spark 3.x, this was &lt;code&gt;false&lt;/code&gt;. That single boolean changes the error semantics of every arithmetic operation, every type cast, and every array access in your codebase.&lt;/p&gt;

&lt;p&gt;What used to return NULL now throws an exception.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;CAST('a' AS INT)&lt;/code&gt;? In 3.x, you got NULL. In 4.x, you get &lt;code&gt;[CAST_INVALID_INPUT]&lt;/code&gt;. &lt;code&gt;2147483647 + 1&lt;/code&gt;? In 3.x, it silently wrapped to a negative number. In 4.x, &lt;code&gt;[ARITHMETIC_OVERFLOW]&lt;/code&gt;. Division by zero, out-of-range array access; same story across the board.&lt;/p&gt;

&lt;p&gt;This is the category of change that passes every smoke test and blows up on real data. Your CSV ingestion pipeline that casts dirty string fields to integers? Dead. Your financial rollup that divides by a denominator that's occasionally zero and falls back to NULL? Dead.&lt;/p&gt;

&lt;p&gt;Databricks called this &lt;a href="https://www.databricks.com/blog/introducing-apache-spark-40" rel="noopener noreferrer"&gt;"one of the most significant shifts in Spark 4.0."&lt;/a&gt; They're right, but "significant shift" undersells it. It's a runtime behavior change with no compile-time warning and no deprecation notice in 3.x. You find out when the job fails at 2am and the on-call page wakes up someone who has no idea what ANSI mode even is.&lt;/p&gt;

&lt;p&gt;The escape hatch exists: set &lt;code&gt;spark.sql.ansi.enabled=false&lt;/code&gt; and keep the old behavior. But that's a stopgap, not a strategy. The Pandas API on Spark already defaults to &lt;code&gt;compute.ansi_mode_support=True&lt;/code&gt; in 4.1+. The window for globally disabling ANSI mode is narrowing with every minor release.&lt;/p&gt;

&lt;p&gt;The real fix is surgical. Spark 4.x ships &lt;code&gt;try_add()&lt;/code&gt;, &lt;code&gt;try_multiply()&lt;/code&gt;, &lt;code&gt;try_cast()&lt;/code&gt;. Per-expression functions that return NULL instead of throwing. Audit your pipelines for every implicit cast and every arithmetic operation that could overflow. Wrap the ones that need NULL semantics in the &lt;code&gt;try_*&lt;/code&gt; variant. Leave everything else strict. You want the errors; they're telling you about data quality problems you've been silently shipping for years.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;ANSI mode doesn't break your pipelines. It reveals the bugs your pipelines have been hiding.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Java 17, the Dependency Audit, and Everything Else That Moved
&lt;/h2&gt;

&lt;p&gt;Java 17 is the minimum runtime for Spark 4.x. &lt;a href="https://spark.apache.org/releases/spark-release-4-0-0.html" rel="noopener noreferrer"&gt;JDK 8 and 11 support dropped entirely.&lt;/a&gt; This one's visible in CI; your build fails, you catch it early. The harder problem is everything downstream.&lt;/p&gt;

&lt;p&gt;The dependency versions jumped across the board. Guava 14 to 33. Jackson 2.15 to 2.18. Hadoop 3.3.4 to 3.4.1. Arrow 12.0.1 to 18.1.0. If you're shading any of these into a fat JAR (and you probably are), your shading rules need a full audit. A library using Guava's &lt;code&gt;Multimap&lt;/code&gt; in version 14 has a different serialized form than version 33. Deserializing old data with new classes can silently succeed with wrong behavior. I spent a full day once debugging a deserialization issue that traced back to exactly this kind of version skew in a shaded dependency. That's silent data corruption; the worst category of bug, and the one you don't find until someone at finance asks why the board deck numbers don't add up.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;javax.servlet&lt;/code&gt; to &lt;code&gt;jakarta.servlet&lt;/code&gt; migration is mandatory. Any custom REST endpoints or transitive dependencies that import &lt;code&gt;javax.servlet&lt;/code&gt; fail at runtime. Scala 2.12 is gone; you're recompiling against 2.13 with its overhauled collections API. Scalafix handles most of it mechanically, but "most" and "all" are different words.&lt;/p&gt;

&lt;p&gt;Then there's Java 17's module system. Spark's NIO code touches JDK internals, which means &lt;code&gt;, add-opens&lt;/code&gt; flags for &lt;code&gt;java.base/sun.nio.ch&lt;/code&gt;, &lt;code&gt;java.base/jdk.internal.misc&lt;/code&gt;, and a growing list of others. I've seen production configs with 6 &lt;code&gt;, add-opens&lt;/code&gt; flags just to keep module-access warnings quiet. It works, but it's configuration debt you'll carry forward to every deploy.&lt;/p&gt;

&lt;p&gt;The ecosystem isn't fully ready either. OpenSearch Hadoop's Spark 4.0 support &lt;a href="https://github.com/opensearch-project/opensearch-hadoop/issues/668" rel="noopener noreferrer"&gt;issue&lt;/a&gt; was filed after launch and remains open. If you depend on connectors outside the core ecosystem, check compatibility before you start. Finding out your connector is blocked after you've done all the Java 17 work is a special kind of frustrating.&lt;/p&gt;

&lt;p&gt;The upside: what shipped alongside these breaking changes is genuinely useful. The VARIANT type (GA in 4.1) gives you native semi-structured data with shredding optimization. SQL UDFs are transparent to the Catalyst optimizer. Pipe syntax (&lt;code&gt;|&amp;gt;&lt;/code&gt;) lets you write SQL that reads top-to-bottom instead of inside-out. Arrow-optimized Python UDFs are on by default in 4.2, delivering faster execution with zero code changes. The upgrade has a real cost, but it's buying you real capabilities.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;PySpark&lt;/strong&gt; at 1.5 MB: What Spark Connect Actually Changes for &lt;strong&gt;Data Engineering&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The new &lt;code&gt;pyspark-client&lt;/code&gt; package is &lt;a href="https://pypi.org/project/pyspark-client/" rel="noopener noreferrer"&gt;1.5 MB on PyPI&lt;/a&gt;. No JVM. No JARs. Pure Python. &lt;code&gt;pip install pyspark-connect&lt;/code&gt; and you're talking to a remote Spark cluster over gRPC.&lt;/p&gt;

&lt;p&gt;For 15 years, &lt;strong&gt;PySpark&lt;/strong&gt; meant bundling a local JVM. Your CI container, your notebook environment, your local dev setup; all needed Java installed. Spark Connect removes that constraint for any workflow running against a remote cluster. AWS launched Spark Connect on EMR Serverless in June 2026. Databricks Runtime 19 GA'd 2 days after Spark 4.2. The platform vendors are converging hard on this architecture.&lt;/p&gt;

&lt;p&gt;The tradeoffs matter, though. Spark Connect sends unresolved logical plans to the server. Schema analysis happens at execution time, not at DataFrame construction. In classic PySpark, &lt;code&gt;df.select("nonexistent_column")&lt;/code&gt; fails immediately. In Spark Connect, it fails at &lt;code&gt;.show()&lt;/code&gt; or &lt;code&gt;.collect()&lt;/code&gt;. &lt;a href="https://learn.microsoft.com/en-us/azure/databricks/spark/connect-vs-classic" rel="noopener noreferrer"&gt;Microsoft's Azure Databricks documentation&lt;/a&gt; spells this out with code examples. If your test suite catches column-name typos at plan construction, those tests now silently pass when they shouldn't. That's a foot-gun for teams migrating without careful integration testing.&lt;/p&gt;

&lt;p&gt;The bigger constraint: RDD operations don't work through Spark Connect. &lt;code&gt;SparkContext&lt;/code&gt;, &lt;code&gt;RDD&lt;/code&gt;, anything touching private JVM methods; none of it crosses the gRPC boundary. Legacy RDD pipelines either get ported to DataFrames or stay on classic PySpark. There's no halfway option.&lt;/p&gt;

&lt;p&gt;And production stability is still rough. Practitioner reports describe Connect servers needing daily preventive restarts and struggling with long-running, resource-intensive jobs that destabilize the shared server. The 1.5 MB client is elegant; the server side needs more operational hardening before you trust it with your most critical batch workloads.&lt;/p&gt;

&lt;p&gt;For greenfield projects, Spark Connect is the clear direction. For existing codebases with RDD usage, this migration has a higher cost than the 1.5 MB headline suggests.&lt;/p&gt;

&lt;h2&gt;
  
  
  The &lt;strong&gt;Spark 4&lt;/strong&gt; Support Window Is Closing
&lt;/h2&gt;

&lt;p&gt;Spark 3.5's extended LTS &lt;a href="https://spark.apache.org/versioning-policy.html" rel="noopener noreferrer"&gt;runs through November 2027&lt;/a&gt;, but "extended LTS" means security fixes only. No features, no performance backports, no bug fixes that aren't CVEs. The Apache Foundation called this extension out explicitly as pressure relief for migration. It's a deadline with a grace period, not indefinite support.&lt;/p&gt;

&lt;p&gt;The managed-service timelines make planning harder. AWS EMR 7.5 (Spark 3.5.6) hits end of standard support November 2026. Databricks Runtime 15.4 LTS carries through August 2027. Google Dataproc hasn't published a Spark 3.5 EOL date at all. Multi-cloud shops can't coordinate a single cutover; you're staggering upgrades per platform or running mixed major versions in parallel. Both cost money and engineer time that nobody budgeted for.&lt;/p&gt;

&lt;p&gt;Here's the career angle. Spark 4 literacy is becoming a baseline expectation in &lt;strong&gt;data engineering&lt;/strong&gt; interviews. ANSI mode semantics, Spark Connect architecture, the &lt;code&gt;try_*&lt;/code&gt; function family; these are the kinds of questions that separate candidates who've done the work from candidates who've read the bullet points. If you're prepping right now, understanding what breaks in the 3.x to 4.x transition matters more than memorizing another API call. Concepts transfer; syntax doesn't. But the concepts here are specific: why does ANSI mode throw on overflow, what does lazy schema analysis mean for testing, when do you reach for &lt;code&gt;try_cast()&lt;/code&gt; vs. fixing the upstream data. That kind of spark practice is what DataDriven is good for, and we designed the prep around exactly these conceptual shifts because they compound everywhere you look.&lt;/p&gt;

&lt;p&gt;The tools change every 18 months. The problems don't. Schema drift, type coercion bugs, upstream teams breaking contracts without telling you. Spark 4.x just made some of those problems louder. I'll take loud failures over silent ones every single time.&lt;/p&gt;

&lt;p&gt;What's the gnarliest thing ANSI mode broke in your pipelines?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>python</category>
      <category>database</category>
      <category>devops</category>
    </item>
    <item>
      <title>dbt Now Runs on Flink. What Changes for Data Engineers</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Thu, 17 Sep 2026 10:10:17 +0000</pubDate>
      <link>https://dev.to/datadriven/dbt-now-runs-on-flink-what-changes-for-data-engineers-55hg</link>
      <guid>https://dev.to/datadriven/dbt-now-runs-on-flink-what-changes-for-data-engineers-55hg</guid>
      <description>&lt;p&gt;For 5 years I maintained 2 separate sets of transformation logic for the same business domain. Batch models in dbt on Snowflake for the warehouse. A completely separate Flink application in Java for the streaming layer. Same business rules, different languages, different testing frameworks, different deploy processes, different on-call runbooks. When something broke on the streaming side, the batch team shrugged. When the warehouse was wrong, the streaming team pointed at their own dashboards and said "looks fine to us."&lt;/p&gt;

&lt;p&gt;That split has defined &lt;strong&gt;data engineering&lt;/strong&gt; careers for the better part of a decade. And it just got smaller.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Confluent&lt;/strong&gt; shipped a free, open-source &lt;strong&gt;dbt&lt;/strong&gt; adapter for &lt;strong&gt;Apache Flink&lt;/strong&gt; &lt;strong&gt;SQL&lt;/strong&gt; in Q2 2026. A parallel community adapter from Xebia covers self-managed Flink. Together, they let you write &lt;strong&gt;streaming&lt;/strong&gt; transformations inside the same dbt project you already use for batch. Same &lt;code&gt;ref()&lt;/code&gt;, same DAG, same &lt;code&gt;dbt test&lt;/code&gt;. Different engine underneath.&lt;/p&gt;

&lt;p&gt;This matters more for your career than for your architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the dbt Adapter for Apache Flink SQL Actually Does
&lt;/h2&gt;

&lt;p&gt;The core idea is straightforward: you write a dbt model in SQL, and instead of compiling to Snowflake or BigQuery, it compiles to Flink SQL and runs as a continuous query against Kafka topics.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.confluent.io/blog/meeting-data-and-analytics-engineers-where-they-are-introducing-the-dbt-adapter-for-confluent-cloud/" rel="noopener noreferrer"&gt;Confluent adapter&lt;/a&gt; supports 3 materialization types. &lt;code&gt;view&lt;/code&gt; creates a virtual table over a Kafka topic. &lt;code&gt;streaming_table&lt;/code&gt; runs a continuous &lt;code&gt;INSERT INTO...SELECT&lt;/code&gt; that writes results to a new topic with changelog semantics. &lt;code&gt;streaming_source&lt;/code&gt; defines a connector-backed source; basically your Kafka topic declaration.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ref()&lt;/code&gt; works identically to batch dbt. The adapter resolves dependencies, builds the DAG, and deploys models in topological order so upstream tables exist before downstream consumers start reading. You can run &lt;code&gt;dbt ls , select model_name+&lt;/code&gt; for selective deployments. If you've used dbt on any warehouse, the workflow is familiar.&lt;/p&gt;

&lt;p&gt;The genuinely clever piece is testing. When a query runs continuously and results are unbounded, how do you write a test with a deterministic pass/fail? You can't just &lt;code&gt;SELECT COUNT(*)&lt;/code&gt; from an infinite stream and expect a stable answer. Confluent's adapter automatically switches to bounded execution mode during tests, using snapshot queries that read Kafka topics up to the current timestamp and return a finite result set. Unit tests use mock data on temporary tables. Your &lt;code&gt;dbt test&lt;/code&gt; command works. It just works differently under the hood.&lt;/p&gt;

&lt;p&gt;That said, data quality tests against live production streams are marked "coming soon" as of the &lt;a href="https://www.confluent.io/blog/2026-q2-confluent-cloud-launch/" rel="noopener noreferrer"&gt;Q2 2026 release&lt;/a&gt;. Unit tests with mocked inputs, yes. Asserting &lt;code&gt;not_null&lt;/code&gt; on a stream that's been running for 3 weeks in prod, not yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Career Cost of 2 Toolchains
&lt;/h2&gt;

&lt;p&gt;I've been on both sides of this hiring wall. I've interviewed candidates who were excellent batch engineers and watched them stumble the moment someone asked about watermarks. I've also interviewed streaming specialists who couldn't design a slowly changing dimension to save their lives. The industry created 2 career tracks for work that, conceptually, is the same discipline.&lt;/p&gt;

&lt;p&gt;The numbers back this up. Data engineers with Flink or Kafka production experience &lt;a href="https://www.ziprecruiter.com/Jobs/Data-Engineer-Flink" rel="noopener noreferrer"&gt;earn a 15-25% salary premium&lt;/a&gt; over equivalent batch-focused roles, with medians around $130k. That premium exists because the supply is thin; streaming has been genuinely harder to hire for because the toolchain was completely separate. Different orchestration, different monitoring, different failure modes, different interview prep.&lt;/p&gt;

&lt;p&gt;67% of enterprises now operate both batch and streaming pipelines. The demand side is exploding. But until this year, becoming a "streaming data engineer" meant learning an entirely separate stack on top of your existing one. Kafka, Flink APIs, custom deployment scripts, bespoke testing harnesses. 2 to 3 months of ramp-up time after you already knew batch fundamentals.&lt;/p&gt;

&lt;p&gt;dbt extending to Flink compresses that ramp. You still need to understand watermarks, windowing, state TTL, and event-time semantics. Those are concepts, and concepts are what actually matter. But you don't need to learn a new transformation framework, a new testing approach, a new CI/CD pipeline, and a new way to express lineage. You use the dbt workflow you already know and learn the streaming concepts on top of it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Junior engineers worry about which tool to learn. Senior engineers worry about which problems to solve. Staff engineers worry about which problems to prevent. The dbt-Flink adapter is a staff-level move: it prevents the problem of maintaining 2 divergent transformation stacks.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This has interview implications too. Watermark literacy is now expected at the junior level in loops that touch streaming. 5 years ago, watermarks were specialist knowledge. In 2026, with 60% of new pipelines carrying real-time components, interviewers expect you to explain &lt;code&gt;WATERMARK FOR order_time AS order_time - INTERVAL '5' SECOND&lt;/code&gt; and why you chose that lateness bound. The concepts haven't changed; the percentage of roles that require them has. If you want reps on windowed aggregations before your next loop, for sql interview prep, try datadriven.io where we've been expanding the streaming SQL problems to match what panels actually ask now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Managed vs. Self-Hosted: 2 Adapters, 2 Trade-offs for Data Engineers
&lt;/h2&gt;

&lt;p&gt;There are actually 2 dbt-Flink adapters, and which one you care about depends on your infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Confluent's adapter&lt;/strong&gt; targets Confluent Cloud exclusively. It ships with a &lt;code&gt;confluent-sql&lt;/code&gt; Python driver (DB-API v2 compliant) that talks directly to Confluent's REST API. The upside is zero Flink cluster management; Confluent handles compute, scaling, and checkpointing. The downside is a hard lock to Kafka as your only data source. If your events live in DynamoDB, S3, or a JDBC database, they need to land in a Kafka topic first. Confluent Cloud Flink pricing runs $0.21 per CFU-hour, billed by the minute.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Xebia's adapter&lt;/strong&gt; (originally from GetInData, released in 2023) targets self-managed Flink via the SQL Gateway. It exposes the full Flink connector ecosystem: Kafka, Elasticsearch, JDBC, S3, HDFS, Kinesis, and 50+ community connectors. You get more flexibility in exchange for operating your own Flink cluster. The trade-off is real; version 1.3.11 added session cluster lifecycle management, but the Xebia team themselves acknowledge the architecture is "lightweight and stateless" at the cost of robustness. If deployment fails between internal steps, you can lose job progress. Savepoint recovery is manual.&lt;/p&gt;

&lt;p&gt;102 GitHub stars on the Xebia adapter after 3 years tells you something about adoption. Self-managed Flink is operationally expensive, and most teams that want streaming are gravitating toward managed offerings. But if you're already running Flink clusters and want dbt integration without a Confluent Cloud dependency, it's the only option available.&lt;/p&gt;

&lt;p&gt;Both adapters implement bounded execution for tests. Both support &lt;code&gt;ref()&lt;/code&gt; and DAG resolution. Both leave data quality tests on live streams incomplete. The split comes down to infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the Maturity Gaps in Data Engineering Tooling Still Bite
&lt;/h2&gt;

&lt;p&gt;I'd be lying if I said this was production-ready the way dbt on Snowflake is production-ready. The gaps are real, and they mostly stem from Flink SQL's architecture rather than adapter immaturity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No snapshots.&lt;/strong&gt; Flink SQL doesn't support &lt;code&gt;MERGE&lt;/code&gt; or the CTE-based updates that dbt snapshots require for Type-2 slowly changing dimensions. If your warehouse workflow depends on &lt;code&gt;dbt snapshot&lt;/code&gt; for audit trails and historical dimension tracking, that pattern doesn't exist here. You'll need a separate solution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No incremental materialization.&lt;/strong&gt; dbt's batch-incremental model (process only new rows since last run) doesn't map to continuous processing. You use &lt;code&gt;streaming_table&lt;/code&gt; instead, which is a fundamentally different execution model. Existing incremental pipelines need a rewrite, not a config change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No cluster provisioning.&lt;/strong&gt; The adapter can't create or drop Kafka clusters. Both &lt;code&gt;dbname&lt;/code&gt; and &lt;code&gt;schema&lt;/code&gt; must reference pre-existing infrastructure. If you're used to Snowflake's fully declarative setup where dbt can create schemas and databases, that gap is noticeable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;10-minute session expiry on Xebia's adapter.&lt;/strong&gt; Tables created in a Flink SQL Gateway session expire after 10 minutes by default. Run &lt;code&gt;dbt test&lt;/code&gt; 15 minutes after &lt;code&gt;dbt run&lt;/code&gt; and you get table-not-found errors. This kind of brittleness doesn't exist in warehouse-based dbt.&lt;/p&gt;

&lt;p&gt;And the conceptual gap remains even when the workflow gap closes. Stateful operations in continuous streams require you to think about state TTL, changelog semantics, and event time vs. processing time ordering. These are streaming-specific concerns that batch engineers have to learn regardless of whether the command they type is still &lt;code&gt;dbt run&lt;/code&gt;. The adapter unifies the workflow; it doesn't unify the knowledge.&lt;/p&gt;

&lt;p&gt;As Kai Waehner at Confluent &lt;a href="https://www.kai-waehner.de/blog/2026/03/26/dbt-meets-apache-flink-one-workflow-for-data-engineers-on-snowflake-bigquery-databricks-and-confluent/" rel="noopener noreferrer"&gt;put it&lt;/a&gt;: "Teams should expect to work with an evolving ecosystem." That's a diplomatic way of saying it's early.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Changes
&lt;/h2&gt;

&lt;p&gt;Here's my honest read. dbt running on Apache Flink is a big deal for the data engineering career path and a modest deal for architecture. Most companies still don't need streaming. Batch handles 90% of what organizations actually run. That hasn't changed.&lt;/p&gt;

&lt;p&gt;What has changed is the cost of adding the other 10%. Previously, streaming meant a separate team, a separate stack, and a 6-figure hiring premium. Now it means learning Flink SQL concepts while keeping the dbt workflow your team already operates. The marginal cost of streaming capability dropped significantly in a single product cycle.&lt;/p&gt;

&lt;p&gt;The Fivetran acquisition of dbt Labs (closed June 2026, combined ~$600M ARR) makes this even more interesting. The merged entity controls both ELT ingestion and SQL transformation. Adding Flink support means they're positioning to own batch and streaming transformation in one platform. That kind of consolidation has historically been great for adoption and terrible for pricing leverage, but that's a problem for next year.&lt;/p&gt;

&lt;p&gt;For now: learn watermarks. Learn windowing. Learn event-time semantics. These are concepts that transfer across any streaming engine. The dbt adapter means you can practice them without rebuilding your entire workflow from scratch. And if you're hiring, stop splitting your DE reqs into "batch" and "streaming" tracks. The toolchain just told you those are converging.&lt;/p&gt;

&lt;p&gt;What's your team's plan? Adopting the adapter, waiting for maturity, or still running separate batch and streaming stacks? I'm curious where people are actually landing on this.&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>sql</category>
      <category>database</category>
      <category>career</category>
    </item>
    <item>
      <title>5,800 Data Engineers on What Is Actually in Production in 2026</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Wed, 16 Sep 2026 18:01:21 +0000</pubDate>
      <link>https://dev.to/datadriven/5800-data-engineers-on-what-is-actually-in-production-in-2026-1570</link>
      <guid>https://dev.to/datadriven/5800-data-engineers-on-what-is-actually-in-production-in-2026-1570</guid>
      <description>&lt;p&gt;Every year, someone publishes a "State of Data Engineering" survey with 200 respondents, mostly from their own Slack community, and calls it representative. I ignore most of them. When &lt;a href="https://www.astronomer.io/press-releases/astronomer-releases-state-of-apache-airflow-2026-report/" rel="noopener noreferrer"&gt;Astronomer dropped their State of Apache Airflow 2026 report&lt;/a&gt; with 5,818 respondents across 122 countries, I paid attention. That's actual sample size.&lt;/p&gt;

&lt;p&gt;I've spent years on both sides of the &lt;strong&gt;data engineering&lt;/strong&gt; interview table, and I've watched candidates prep for tools that their target companies don't even use. The numbers in this &lt;strong&gt;survey&lt;/strong&gt; confirm a few things I've seen firsthand, contradict some of the loudest takes on LinkedIn, and surface a gap that should make every data team uncomfortable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What 5,800 Practitioners Are Actually Running in Production
&lt;/h2&gt;

&lt;p&gt;The survey ran September 15 to November 20, 2025. 50 questions. 5,818 practitioners. 122 countries. Astronomer calls it the largest data engineering survey ever conducted. The &lt;a href="https://airflow.apache.org/blog/airflow-survey-2025/" rel="noopener noreferrer"&gt;2025 edition had 5,250 respondents from 116 countries&lt;/a&gt;, so this is roughly 10% growth year over year. Steady accumulation.&lt;/p&gt;

&lt;p&gt;One number jumped out immediately: &lt;strong&gt;Apache Airflow&lt;/strong&gt; now has over 3,600 unique contributors. More than Spark's 2,000+. More than Kafka's 1,530. That's 135% more contributors than Kafka and 80% more than Spark. 3,600 humans who committed code to the project.&lt;/p&gt;

&lt;p&gt;The contributor gap tells you something about the &lt;strong&gt;orchestration&lt;/strong&gt; layer specifically. Airflow's ecosystem of provider packages (cloud connectors, custom operators, executor integrations) has lowered the friction for first-time contributors in a way that Spark's heavier core or Kafka's narrower domain focus hasn't matched. Contributor velocity is the best leading indicator of whether an open-source project will keep pace with the ecosystem around it. Right now, Airflow is outpacing both of its most obvious peers, and that matters more than any vendor's marketing slide.&lt;/p&gt;

&lt;h2&gt;
  
  
  84% Plan the Airflow 3 Migration. 26% Have Done It.
&lt;/h2&gt;

&lt;p&gt;This gap is the most telling number in the entire report.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Airflow 3&lt;/strong&gt; shipped in April 2025. Less than a year later, 26% of respondents have completed the migration. 84% of those still on 2.x say they're planning it. That 58-point spread between planning and doing reflects engineering teams staring at a list of breaking changes and budgeting real project time.&lt;/p&gt;

&lt;p&gt;The breaking changes aren't cosmetic. Airflow 3 removes SubDAGs entirely. It eliminates context variables like &lt;code&gt;execution_date&lt;/code&gt;, &lt;code&gt;prev_ds&lt;/code&gt;, and &lt;code&gt;next_ds&lt;/code&gt;. Workers can no longer directly access the metadata database; everything routes through the REST API. &lt;code&gt;xcom_pull(key=)&lt;/code&gt; stops searching upstream tasks. Each of these forces full DAG rewrites. If you've got 300 DAGs in &lt;strong&gt;production&lt;/strong&gt; and half of them use &lt;code&gt;execution_date&lt;/code&gt; for partition logic, you're looking at weeks of refactoring before you even touch your CI pipeline. That's why the gap exists.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.astronomer.io/blog/state-of-airflow-2026/" rel="noopener noreferrer"&gt;Astronomer's own customer base&lt;/a&gt; tells a different story: 48% already run Airflow 3, and among their largest enterprise customers (50,000+ employees), 60% have deployed it. The gap between managed-platform customers and the broader community is a resource story. These companies have dedicated platform teams that can absorb migration work while the rest of the org keeps shipping. The 4-person data team at a Series B startup running 150 DAGs? They're going to plan the migration for Q3 and push it to Q4. Then Q1.&lt;/p&gt;

&lt;p&gt;And the clock is ticking. &lt;a href="https://kestra.io/blogs/2026-04-06-airflow-2-end-of-life" rel="noopener noreferrer"&gt;Airflow 2 hit end of life in April 2026&lt;/a&gt;. No more security patches. No more bug fixes. No more provider package updates. If you're running Snowflake, Databricks, or BigQuery providers on Airflow 2, you're on borrowed time. Those vendors will start dropping 2.x compatibility in their own provider packages, and that second wave of pressure arrives 6 to 12 months after the official EOL. SOC 2, HIPAA, PCI-DSS; any certification that requires supported software makes this a compliance conversation on top of a technical one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;84% of teams know they need to migrate. 26% have done it. The gap is the actual cost of rewriting production DAGs while keeping the lights on.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you're interviewing right now, expect questions on both Airflow 2 and 3 for at least another 12 months. The industry hasn't turned over yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The GenAI Production Gap
&lt;/h2&gt;

&lt;p&gt;Here's where the &lt;strong&gt;survey&lt;/strong&gt; gets uncomfortable.&lt;/p&gt;

&lt;p&gt;Among all &lt;strong&gt;Airflow&lt;/strong&gt; users, 32% report running GenAI or MLOps workloads in production. Among Astronomer's managed-platform customers, 62%. Among organizations that have been Astronomer customers for 2+ years, 83%.&lt;/p&gt;

&lt;p&gt;51 percentage points between the general population and long-tenured platform customers. Read that again.&lt;/p&gt;

&lt;p&gt;This has nothing to do with ambition. It has everything to do with infrastructure maturity compounding over time. Teams that spent years building out observability, error handling, schema governance, and pipeline idempotency can bolt on GenAI workloads because the plumbing already exists. Teams still figuring out how to make their batch jobs reliable aren't ready to orchestrate inference pipelines; they know it, and they're right to wait.&lt;/p&gt;

&lt;p&gt;Separate research from &lt;a href="https://www.k2view.com/genai-adoption-survey-2026" rel="noopener noreferrer"&gt;K2view's 2026 enterprise survey&lt;/a&gt; backs this up: 76% of organizations cite data quality and consistency as their top barrier to production GenAI, and 62% say their enterprise data simply isn't ready. The bottleneck is the data platform underneath the model. MIT research referenced in the same analysis found 95% of enterprise AI pilots deliver zero P&amp;amp;L impact, which tells you the pilot-to-production gap is structural. You can't buy your way past it with a fancier model.&lt;/p&gt;

&lt;p&gt;Meanwhile, 80%+ of the Astronomer survey respondents say they use AI tools to write pipelines but report that those tools hallucinate, lack context, and generate outdated syntax. I've been saying this for years: AI is making coding interviews increasingly pointless as a signal of engineering ability, but AI as a pipeline developer is nowhere close to replacing the engineer who knows why the pipeline broke at 3am last Tuesday and how to prevent it from happening again.&lt;/p&gt;

&lt;p&gt;89% of Airflow users expect to expand &lt;strong&gt;orchestration&lt;/strong&gt; into revenue-generating, external-facing solutions in 2026. The orchestration layer is moving from back-office utility to core product infrastructure. Your job as a data engineer just got a lot more visible to the business. That's both good and terrifying.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Platform Split and What It Means for Your Next Data Engineering Interview
&lt;/h2&gt;

&lt;p&gt;The conventional wisdom on LinkedIn says Databricks is running away with the market. The survey says: slow down.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Snowflake&lt;/strong&gt;: 36.6%. &lt;strong&gt;Databricks&lt;/strong&gt;: 34.7%. &lt;strong&gt;BigQuery&lt;/strong&gt;: 27.8%. The total spread between first and third is under 9 percentage points. Snowflake leads Databricks by less than 2. Calling a winner here is irresponsible.&lt;/p&gt;

&lt;p&gt;22% of Airflow users run 2 or more of the 3 major cloud data platforms. 1 in 5 teams operates a heterogeneous stack by choice. Intentional multi-platform architecture, not tech debt.&lt;/p&gt;

&lt;p&gt;For interviews, this matters. A lot.&lt;/p&gt;

&lt;p&gt;Snowflake loops emphasize the 3-layer architecture (storage, compute, cloud services), warehouse cost control, clustering strategies, and Snowpipe ingestion patterns. Databricks loops focus on Spark internals (shuffle semantics, partitioning), Delta Lake ACID transactions, and distributed systems thinking under concurrency pressure. These are fundamentally different prep surfaces. A single "data engineering interview" study plan that doesn't branch by platform will teach you the wrong material for half the roles you apply to.&lt;/p&gt;

&lt;p&gt;Check the job posting. If it says Snowflake, drill Snowflake. If it says Databricks, drill Databricks. If it says both, they probably don't know what they want yet; ask in the recruiter screen. Don't waste 3 weeks studying the wrong platform because some influencer told you "Databricks is winning."&lt;/p&gt;

&lt;p&gt;The concept-over-tool argument still holds, though. Data modeling, query optimization, understanding why things break; these transfer across all 3 platforms. The engineers who understand grain, normal forms, and late-arriving data will pass interviews on any platform after a week of syntax review. The ones who memorized &lt;code&gt;COPY INTO&lt;/code&gt; but can't explain a slowly changing dimension will struggle regardless of which logo shows up in the job description. And if the loop includes a Python round, which most of them do now, python interview questions are what DataDriven is good for; we built the problem sets around the patterns that actually show up in production-focused loops, not LeetCode puzzles divorced from the job.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Survey Doesn't Tell You
&lt;/h2&gt;

&lt;p&gt;Astronomer sponsors this survey and operates the platform it measures. The report blends community responses with "trends from Astro customer usage data," which means the 5,818 figure mixes voluntary respondents with inferred signals from their own product telemetry. Response rate and confidence intervals aren't published. The raw data isn't public. When the company running the survey also sells the managed version of the tool, numbers showing managed customers outperforming on every metric deserve some scrutiny.&lt;/p&gt;

&lt;p&gt;Does that invalidate the findings? No. 5,800+ respondents across 122 countries is still the largest &lt;strong&gt;data engineering&lt;/strong&gt; sample I've seen. The platform split, the migration gap, the GenAI adoption curve; these are directionally useful even with selection bias baked in. Use them as a compass, not a GPS coordinate.&lt;/p&gt;

&lt;p&gt;The tools will keep changing every 18 months, like they always have. The problems won't. Schema drift, late-arriving data, upstream teams breaking contracts without telling you. Those are eternal. This survey confirms what I've watched play out for years: the industry is growing, the orchestration layer is becoming more strategic, and the gap between "I'm planning to do this" and "I've actually shipped it" is where careers get made.&lt;/p&gt;

&lt;p&gt;What's your team actually running in production right now? Does it match what the survey says, or are you living in a completely different reality?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>python</category>
      <category>career</category>
      <category>devops</category>
    </item>
    <item>
      <title>Apache Iceberg v3 Is GA. Here Is What Data Engineers Get.</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 15 Sep 2026 10:13:03 +0000</pubDate>
      <link>https://dev.to/datadriven/apache-iceberg-v3-is-ga-here-is-what-data-engineers-get-51ol</link>
      <guid>https://dev.to/datadriven/apache-iceberg-v3-is-ga-here-is-what-data-engineers-get-51ol</guid>
      <description>&lt;p&gt;Every vendor blog post I've read this year about &lt;strong&gt;Apache Iceberg&lt;/strong&gt; v3 says the same thing: faster deletes, upgrade now. &lt;a href="https://docs.snowflake.com/en/release-notes/2026/other/2026-05-07-iceberg-v3-ga" rel="noopener noreferrer"&gt;Snowflake moved v3 to GA on May 7&lt;/a&gt;. Databricks shipped support in Runtime 18.0. Apache Iceberg 1.11.0 landed May 19 and declared the spec production-stable. AWS followed over the summer with Glue 6.0, Redshift, and S3 Tables all claiming v3 readiness.&lt;/p&gt;

&lt;p&gt;The marketing is loud. The substance is real but narrower than the press releases suggest. I've been digging through the actual spec changes, testing engine support matrices, and cataloging the gotchas nobody puts in the blog title. Here's what working &lt;strong&gt;data engineering&lt;/strong&gt; teams actually get, and where the gaps are.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deletion Vectors Solve a File Problem
&lt;/h2&gt;

&lt;p&gt;The headline feature of &lt;strong&gt;Iceberg v3&lt;/strong&gt; is &lt;strong&gt;deletion vectors&lt;/strong&gt;, and the way vendors frame them has been consistently misleading. Every blog leads with "10x faster deletes." The real win is architectural, and understanding the architecture is what separates a surface-level take from a staff-level one.&lt;/p&gt;

&lt;p&gt;In v2, every delete operation produced a separate positional delete file. Delete 500 rows across 1,000 data files and you've spawned 1,000 new delete files in your metadata layer. Run CDC pipelines that touch millions of rows daily, and those delete files multiply fast. Read queries then had to join each data file against its corresponding delete files to figure out which rows were still alive. That join overhead is where read performance died.&lt;/p&gt;

&lt;p&gt;v3 replaces positional delete files with &lt;strong&gt;Roaring bitmaps&lt;/strong&gt; stored in Puffin sidecar files. Each data file pairs 1:1 with a single deletion vector. A transaction deleting rows across 1,000 data files packs all 1,000 deletion vectors into one Puffin file. At read time the engine checks a bitmap, O(1) lookup per row, instead of joining against a pile of delete files.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.dremio.com/blog/dremio-iceberg-v3-deletion-vectors/" rel="noopener noreferrer"&gt;Dremio's testing&lt;/a&gt; showed 50 to 80 percent read improvement when deletion vectors replaced v2 positional deletes. Merge-on-Read workloads under high churn saw 5 to 10x speedups. Real numbers. But they measure something specific: the elimination of join overhead between data files and delete files. A table that rarely sees deletes barely notices the upgrade.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The performance gain from deletion vectors comes from killing file sprawl. If your v2 tables accumulate hundreds of positional delete files per day from CDC, v3 is an immediate quality-of-life upgrade. If they don't, the benefit is real but modest.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The economics reinforce this. Every v2 delete file was an S3 object. Every read that referenced those files issued GET requests. Multiply that across a CDC-heavy lakehouse and the API charges add up fast. v3 collapses all of it into bitmaps. Storage is 2 cents a GB; the S3 GET bill at scale is where the real money hides.&lt;/p&gt;

&lt;p&gt;Compaction is still mandatory. Deletion vectors shrink the metadata cost of tracking dead rows, but the underlying data files those rows live in don't clean themselves up. You still need periodic compaction to reclaim space. The operational discipline hasn't changed; the file-level mechanics got cleaner.&lt;/p&gt;

&lt;h2&gt;
  
  
  Row Lineage: Native CDC at the Format Level
&lt;/h2&gt;

&lt;p&gt;The other big v3 addition is &lt;strong&gt;row lineage&lt;/strong&gt;, and this matters for anyone building change data capture on an &lt;strong&gt;open table format&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;v3 tables carry 2 implicit metadata fields on every row: &lt;code&gt;_row_id&lt;/code&gt; (a unique 64-bit identifier assigned at first write, surviving updates, compactions, and partition changes) and &lt;code&gt;_last_updated_sequence_number&lt;/code&gt; (the snapshot sequence when the row was last modified). The &lt;a href="https://docs.aws.amazon.com/prescriptive-guidance/latest/apache-iceberg-on-aws/table-spec-v3.html" rel="noopener noreferrer"&gt;AWS prescriptive guidance&lt;/a&gt; describes it as "engine-agnostic and interoperable, built into the Iceberg V3 specification, alleviating the need for custom, engine-specific change tracking implementations."&lt;/p&gt;

&lt;p&gt;Before v3, CDC on Iceberg meant full-table diffs or bolting on external tooling to detect changes. I've built those pipelines. They work, in the same way duct tape on a leaking pipe works. Row lineage makes change detection a filter predicate: query rows where &lt;code&gt;_last_updated_sequence_number&lt;/code&gt; exceeds your last checkpoint and you have your delta. No full scans. No external watermark tables.&lt;/p&gt;

&lt;p&gt;The theory is clean. Production is messier.&lt;/p&gt;

&lt;p&gt;After upgrading from v2 to v3, existing rows carry &lt;code&gt;_row_id = null&lt;/code&gt; until they're rewritten via compaction or update. Your CDC pipeline needs null handling during the transition, or you schedule a full compaction pass post-upgrade to backfill lineage across the table. Multiple migration guides call this out explicitly: compaction after a v2-to-v3 upgrade is required for homogeneous lineage coverage.&lt;/p&gt;

&lt;p&gt;The other catch: &lt;code&gt;_row_id&lt;/code&gt; and &lt;code&gt;_last_updated_sequence_number&lt;/code&gt; aren't returned by &lt;code&gt;SELECT *&lt;/code&gt; on all engines. StarRocks and Presto have open GitHub issues where these columns require explicit SELECT or silently drop rows with null lineage values. ClickHouse has a known bug where filtering on lineage columns omits pre-upgrade rows entirely. If your read path runs through anything outside the Spark/Databricks/Snowflake core, test before you depend on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  VARIANT Replaces the JSON String Hack
&lt;/h2&gt;

&lt;p&gt;Every data engineering team I've worked on has jammed semi-structured data into a VARCHAR column as a JSON string at some point. It works. The query performance is awful, the storage efficiency is bad, and whoever has to parse that column downstream curses your name. But it ships.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;VARIANT type&lt;/strong&gt; in v3 stores semi-structured data in binary encoding that preserves native types. A date field inside a VARIANT is typed as a date, not a string that looks like one. &lt;a href="https://www.snowflake.com/en/blog/engineering/apache-iceberg-v3-variant-type/" rel="noopener noreferrer"&gt;Snowflake's benchmarks on 500-million-row tables&lt;/a&gt; showed 31.4% space savings over JSON strings and a 2.3x total query speedup.&lt;/p&gt;

&lt;p&gt;The mechanism behind the performance is &lt;strong&gt;shredding&lt;/strong&gt;: at write time, frequently occurring fields get extracted into typed Parquet columns alongside the binary VARIANT blob. The query engine pushes predicates to those shredded columns, applies row-group and page-level skipping, and only decodes the full VARIANT binary for rows that pass filters. You get the flexibility of schema-on-read with the performance you'd expect from typed columnar storage.&lt;/p&gt;

&lt;p&gt;Shredding carries roughly 35% write overhead. Streaming-heavy or write-dominated workloads might still prefer VARCHAR(JSON) and eat the query cost. The optimization pays off on write-once, read-many patterns. Know your access patterns before flipping the switch.&lt;/p&gt;

&lt;p&gt;Rounding out the v3 spec: default column values, geometry and geography types, nanosecond timestamps, and multi-argument partition transforms. Useful additions. But deletion vectors, row lineage, and VARIANT are the 3 features that change pipeline design.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Platform Write Gap Nobody Advertises
&lt;/h2&gt;

&lt;p&gt;Here's where the distance between announcement and reality gets uncomfortable.&lt;/p&gt;

&lt;p&gt;Snowflake's v3 is GA, but new Snowflake-managed Iceberg tables still default to v2. Opt-in required. In-place v2-to-v3 upgrades? Not supported on Snowflake. You're looking at full table recreation through export and reimport. External engine write support didn't ship until May 26, 19 days after the v3 GA announcement, and it only covers Snowflake-managed tables. That's a narrower story than the press release tells.&lt;/p&gt;

&lt;p&gt;Databricks has the most complete v3 support today. Runtime 18.0 enables deletion vectors and row lineage by default on new v3 tables in Unity Catalog. If you're Databricks-native, the path is clear.&lt;/p&gt;

&lt;p&gt;AWS is fragmented. &lt;a href="https://hidekazu-konishi.com/entry/apache_iceberg_v3_on_aws.html" rel="noopener noreferrer"&gt;Athena can't read v3 tables at all&lt;/a&gt;; the engine throws "Cannot read unsupported version 3." If Athena is your primary analytics engine, v3 blocks 100% of your analytical queries on any table you upgrade. Redshift supports v3 but silently changes timestamp mappings and drops support for VARIANT, STRUCT, LIST, MAP, and GEOMETRY. Glue 6.0 shipped v3 in August, but enabling VARIANT disables Lake Formation fine-grained access control and managed compaction, and it's only available in 15 regions.&lt;/p&gt;

&lt;p&gt;The Python ecosystem is the quiet bottleneck. PyIceberg reads v3 tables but can't write them. ML and data science teams that ingest through Python are stuck on v2 for writes. Open-source Trino doesn't support v3 either; only the proprietary Starburst Galaxy fork does. The spec is converging. The implementations are not. That gap defines the operational reality for the next 12 months.&lt;/p&gt;

&lt;h2&gt;
  
  
  Interviews, Migration, and What to Do About It
&lt;/h2&gt;

&lt;p&gt;v3 internals are already landing in hiring loops. If you're interviewing at companies running lakehouse architectures, expect questions on how deletion vectors differ from positional deletes, how row lineage enables CDC without full scans, and when merge-on-read beats copy-on-write. These are concepts questions, and concepts transfer across engines. Knowing the Databricks-specific API for enabling deletion vectors doesn't tell an interviewer much; understanding &lt;em&gt;why&lt;/em&gt; bitmaps replaced positional deletes does. If you're prepping for these conversations, that's exactly why we made sure data engineering practice problems are covered on datadriven; v3 architecture is the kind of topic where understanding the mechanics matters more than memorizing syntax.&lt;/p&gt;

&lt;p&gt;My actual advice on migration: don't rush. If v2 is stable and your workloads are healthy, v3 is worth planning for on your own schedule. Map your engine dependencies first. If Athena or PyIceberg sits in your critical path, you're blocked until those engines catch up. If you're Databricks-native or running Spark on EMR, start migrating non-critical tables and see how compaction, lineage coverage, and downstream consumers behave before you touch anything finance depends on.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;open table format&lt;/strong&gt; convergence story looks great in keynotes. The engine support matrix tells a messier story. Apache Iceberg v3 delivers real improvements to how lakehouse architectures handle deletes, CDC, and semi-structured data. The spec earned its GA label. Whether your entire stack can use it today is a separate question, and the answer depends more on your engine mix than on the spec itself.&lt;/p&gt;

&lt;p&gt;What's your team's v3 timeline? Blocked on engine support, running it in production, or still on v2 and not feeling any pressure to move?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>database</category>
      <category>sql</category>
      <category>career</category>
    </item>
    <item>
      <title>Fivetran and dbt Merged. Here Is What Data Engineers Get.</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 08 Sep 2026 10:11:59 +0000</pubDate>
      <link>https://dev.to/datadriven/fivetran-and-dbt-merged-here-is-what-data-engineers-get-27e2</link>
      <guid>https://dev.to/datadriven/fivetran-and-dbt-merged-here-is-what-data-engineers-get-27e2</guid>
      <description>&lt;p&gt;I've been through 3 mergers as a practitioner on the receiving end. Every time, the press release says "nothing changes for customers." Every time, something changes for customers.&lt;/p&gt;

&lt;p&gt;On June 1, 2026, &lt;strong&gt;Fivetran&lt;/strong&gt; and &lt;strong&gt;dbt Labs&lt;/strong&gt; &lt;a href="https://www.fivetran.com/press/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents" rel="noopener noreferrer"&gt;completed their all-stock merger&lt;/a&gt;, combining the dominant ingestion platform and the dominant transformation framework under one roof. The same day, &lt;strong&gt;dbt Core&lt;/strong&gt; v2.0 shipped in alpha. The combined company serves more than 100,000 data teams and approaches $600 million in ARR.&lt;/p&gt;

&lt;p&gt;One company now owns both layers of the modern data stack that most of you depend on. Let's talk about what you actually get.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Merger Means for Data Engineering Teams
&lt;/h2&gt;

&lt;p&gt;George Fraser stays CEO. Tristan Handy becomes President. The deal is all-stock, no cash component disclosed. The combined entity counts OpenAI, Zendesk, Coupa, and HubSpot among its named customers.&lt;/p&gt;

&lt;p&gt;Here's the context most coverage glosses over. Fivetran raised $565M at a &lt;a href="https://finance.yahoo.com/news/data-integration-startup-fivetran-raises-153415498.html" rel="noopener noreferrer"&gt;$5.6B valuation in September 2021&lt;/a&gt;. dbt Labs raised $222M at a &lt;a href="https://www.nasdaq.com/press-release/dbt-labs-raises-$222m-in-series-d-funding-at-$4-2b-valuation-led-by-altimeter-with-participation-from-databricks-and-snowflake" rel="noopener noreferrer"&gt;$4.2B valuation in February 2022&lt;/a&gt;. Both valuations were peak-ZIRP numbers. An all-stock merger in 2026 means those numbers got renegotiated by reality.&lt;/p&gt;

&lt;p&gt;The strategic logic is straightforward: 80 to 90% of Fivetran's customer base already runs dbt. As &lt;a href="https://datacoves.com/post/dbt-fivetran" rel="noopener noreferrer"&gt;Datacoves noted&lt;/a&gt;, the architectural lock-in was already present; the commercial incentive to make the pair inseparable is now explicit. Fivetran's recent pricing moves reinforce the skepticism. Starting January 2026, Fivetran shifted to per-connection billing with $5 minimums and began &lt;a href="https://www.fivetran.com/pricing" rel="noopener noreferrer"&gt;counting deletes toward MAR&lt;/a&gt;, bumping typical small-team costs from roughly $500/month to $1,200 to $1,800. That history makes the "we'll keep dbt Core open" pledge read differently than it would from a vendor with no recent price shock.&lt;/p&gt;

&lt;p&gt;dbt's strength before this deal was tool-agnosticism. You could pair it with Fivetran, Airbyte, Stitch, or a pile of custom loaders. That neutrality is now a question mark. Not because anyone announced a restriction, but because the incentive structure shifted overnight.&lt;/p&gt;

&lt;h2&gt;
  
  
  dbt Core v2.0: The Rust Rewrite
&lt;/h2&gt;

&lt;p&gt;The headline feature of v2.0 is the Rust-based Fusion engine, replacing the Python runtime. dbt Labs claims &lt;a href="https://docs.getdbt.com/blog/dbt-fusion-engine" rel="noopener noreferrer"&gt;up to 30x faster parsing&lt;/a&gt; on projects with 10,000+ models and 2x faster full-project compilation overall. A 5,000-model project reportedly parses in under 4 seconds on a 16-core workstation.&lt;/p&gt;

&lt;p&gt;These numbers are real, but conditional. The 30x figure is for the largest projects. The benchmarks don't disclose hardware specs or methodology for mid-sized projects in the 1,000 to 5,000 model range. If you're running 200 models, your parse time was already fine.&lt;/p&gt;

&lt;p&gt;The Rust rewrite kills the GIL bottleneck, meaning &lt;a href="https://www.getdbt.com/blog/dbt-fusion-experience" rel="noopener noreferrer"&gt;all parsing and compilation happen locally&lt;/a&gt; without needing cloud compute to go fast. That's a genuine win. The v2 parser also enforces strict syntax: undefined macros and missing variables now fail at parse time instead of silently passing through to compile. This will break things. That's the point.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;dbt Core stays Apache 2.0.&lt;/strong&gt; The Fusion runtime that was previously proprietary is now &lt;a href="https://docs.getdbt.com/blog/dbt-core-v2-is-here" rel="noopener noreferrer"&gt;open source&lt;/a&gt;. This is, on paper, the opposite of the license-restriction pattern everyone fears. Credit where it's due: they didn't have to do that.&lt;/p&gt;

&lt;p&gt;But here's the part that matters for your wallet.&lt;/p&gt;

&lt;h3&gt;
  
  
  The 30% Cost Reduction Claim
&lt;/h3&gt;

&lt;p&gt;The "30%+ infrastructure cost reduction" you keep seeing in headlines refers to &lt;strong&gt;dbt State&lt;/strong&gt;, a caching layer that skips unchanged models. It's a separate, metered service. The Rust engine makes your laptop faster; &lt;a href="https://www.getdbt.com/blog/dbt-state-use-case" rel="noopener noreferrer"&gt;dbt State cuts your warehouse bill&lt;/a&gt;. Different products, different problems.&lt;/p&gt;

&lt;p&gt;dbt State works by caching SQL query hashes and warehouse timestamps, then &lt;a href="https://docs.getdbt.com/docs/deploy/dbt-state-about" rel="noopener noreferrer"&gt;only building models where upstream data or code actually changed&lt;/a&gt;. Obie Insurance &lt;a href="https://www.getdbt.com/blog/how-obie-cut-compute-costs-by-30-percent" rel="noopener noreferrer"&gt;reported 30% compute savings&lt;/a&gt;. EQT hit 50% cost reduction and 60% faster runtimes. dbt Labs internally claims &lt;a href="https://www.getdbt.com/blog/fivetran-dbt-20-future" rel="noopener noreferrer"&gt;64% compute savings, roughly $400,000 annually&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The catch: dbt State bills at &lt;a href="https://docs.getdbt.com/docs/platform/billing/dbt-state-usage" rel="noopener noreferrer"&gt;$0.094 per Daily Active Target Table&lt;/a&gt;. A DATT is any model, test, or seed that was skipped, cloned, or reused on a given day. You need to model the cost at your project's size today, then at 3 years out, and decide whether the warehouse savings still clear the bill. dbt's own documentation &lt;a href="https://docs.getdbt.com/docs/platform/billing/optimize-costs" rel="noopener noreferrer"&gt;warns about exactly this&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Also: Python models always rebuild. dbt State &lt;a href="https://docs.getdbt.com/docs/deploy/dbt-state-about" rel="noopener noreferrer"&gt;cannot reuse them&lt;/a&gt;. If your team leans on Python for ML feature engineering, that layer gets zero benefit from the caching story.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The tools change every 18 months. The problems don't. Schema drift, late-arriving data, upstream teams breaking contracts without telling you. Those are eternal. A merger doesn't change that calculus.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Open Source Question
&lt;/h2&gt;

&lt;p&gt;I've watched this play out before. HashiCorp &lt;a href="https://spacelift.io/blog/terraform-license-change" rel="noopener noreferrer"&gt;moved Terraform from MPL 2.0 to BSL 1.1&lt;/a&gt; in August 2023. Redis &lt;a href="https://techcrunch.com/2024/03/21/redis-switches-licenses-acquires-speedb-to-go-beyond-its-core-in-memory-database/" rel="noopener noreferrer"&gt;abandoned BSD for SSPL&lt;/a&gt; in March 2024. Both cited cloud providers repackaging their open code. Both communities saw the pattern: &lt;strong&gt;open source&lt;/strong&gt; commitments made during fundraising yield to commercial pressure once the cap table needs returns.&lt;/p&gt;

&lt;p&gt;dbt Labs says they're going in the opposite direction. They &lt;a href="https://docs.getdbt.com/blog/dbt-core-v2-is-here" rel="noopener noreferrer"&gt;open-sourced the Fusion engine&lt;/a&gt; and publicly stated that "&lt;a href="https://www.getdbt.com/blog/fivetran-dbt-20-future" rel="noopener noreferrer"&gt;we have wanted to ship more code in the open, not less&lt;/a&gt;." I believe that's genuine today.&lt;/p&gt;

&lt;p&gt;The risk isn't license restriction. The risk is feature parity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;dbt Wizard&lt;/strong&gt;, the AI-assisted authoring agent, launched in &lt;a href="https://www.getdbt.com/product/dbt-wizard" rel="noopener noreferrer"&gt;public beta on June 1&lt;/a&gt;. It requires a Starter, Enterprise, or Enterprise+ account. It's &lt;a href="https://docs.getdbt.com/docs/platform/wizard-overview" rel="noopener noreferrer"&gt;billed per token&lt;/a&gt; starting September 1, 2026. &lt;strong&gt;dbt Mesh&lt;/strong&gt;, &lt;strong&gt;dbt Copilot&lt;/strong&gt;, and &lt;strong&gt;dbt Canvas&lt;/strong&gt; are all &lt;a href="https://docs.getdbt.com/docs/cloud/enable-dbt-copilot" rel="noopener noreferrer"&gt;Enterprise-tier exclusives&lt;/a&gt; at $200 to $400 per developer seat per month.&lt;/p&gt;

&lt;p&gt;Brooklyn Data Co. put it plainly: "&lt;a href="https://www.brooklyndata.co/ideas/2026/06/01/dbt-core-v2_0-is-here-and-the-real-story-isnt-the-one-in-the-headline" rel="noopener noreferrer"&gt;The tool you're told to reach for is no longer the fully-open one.&lt;/a&gt;" When dbt Labs' own recommended path includes closed-source components, the open source label on Core starts carrying less practical weight.&lt;/p&gt;

&lt;p&gt;The Hacker News crowd &lt;a href="https://news.ycombinator.com/item?id=45568842" rel="noopener noreferrer"&gt;summed up the fear&lt;/a&gt;: "Core is gonna stay unchanged while Cloud keeps gaining new features. Eventually, it will be end-of-life." I wouldn't make that prediction; Databricks continues investing in open-source Spark and Delta Lake while competing commercially. But Nexla's analysis &lt;a href="https://nexla.com/blog/open-source-vs-saas-fivetran-dbt-merger" rel="noopener noreferrer"&gt;correctly notes&lt;/a&gt; that "Fivetran is a company with zero history in open source and every incentive to drive users toward their paid products."&lt;/p&gt;

&lt;p&gt;The merger didn't kill open source dbt. But it placed every high-visibility AI feature behind consumption pricing. The letter of the promise is intact. The spirit is evolving.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Data Engineers Should Do Before Touching v2
&lt;/h2&gt;

&lt;p&gt;The release candidate &lt;a href="https://releasebot.io/updates/dbt-labs/dbt-core" rel="noopener noreferrer"&gt;shipped September 2&lt;/a&gt;. No GA date announced. dbt Labs' &lt;a href="https://github.com/dbt-labs/dbt-core/blob/main/docs/roadmap/2026-06-announcing-v2.md" rel="noopener noreferrer"&gt;roadmap document&lt;/a&gt; explicitly asks for community testing "against ever-bigger DAGs, custom macro implementations, and varied system configurations." Translation: they expect edge cases to surface.&lt;/p&gt;

&lt;p&gt;Here's your checklist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test the v2 parser on v1.&lt;/strong&gt; dbt Core v1.12 includes an &lt;a href="https://docs.getdbt.com/reference/global-configs/parsing" rel="noopener noreferrer"&gt;opt-in &lt;code&gt;, use-v2-parser&lt;/code&gt; flag&lt;/a&gt;. Run &lt;code&gt;dbt parse , use-v2-parser&lt;/code&gt; against your project. If it passes, you're probably fine. If it throws on undefined macros or missing variables, those are real problems the v1 parser was silently ignoring.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Audit your CLI scripts.&lt;/strong&gt; &lt;code&gt;, models&lt;/code&gt; becomes &lt;code&gt;, select&lt;/code&gt;. &lt;code&gt;, partial-parse&lt;/code&gt; is gone entirely. &lt;a href="https://docs.getdbt.com/docs/dbt-versions/dbt-upgrade/upgrading-to-v2" rel="noopener noreferrer"&gt;Every CI/CD pipeline, Makefile, and wrapper script&lt;/a&gt; that invokes dbt needs a pass.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check your artifact consumers.&lt;/strong&gt; v2 produces &lt;a href="https://datapace.ai/blog/dbt-core-2-artifacts-metadata-contract" rel="noopener noreferrer"&gt;Parquet metadata artifacts instead of JSON&lt;/a&gt;. If you have tooling that parses &lt;code&gt;manifest.json&lt;/code&gt; for lineage, docs, or metadata pipelines, it will break. This is the migration risk most teams discover weeks after deployment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verify your adapter.&lt;/strong&gt; Adapters for BigQuery, Databricks, Redshift, Snowflake, Spark, and DuckDB are &lt;a href="https://docs.getdbt.com/docs/dbt-versions/dbt-upgrade/upgrading-to-v2" rel="noopener noreferrer"&gt;in preview or beta&lt;/a&gt;. If you're on something else, check before you upgrade.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't confuse RC with GA.&lt;/strong&gt; A release candidate is pre-production. Run it in a dev environment, compare outputs against v1, and wait for the stable release before migrating production DAGs.&lt;/p&gt;

&lt;p&gt;Meanwhile, the &lt;strong&gt;Agents Schema&lt;/strong&gt; initiative (a &lt;a href="https://github.com/dbt-labs/agents_schema" rel="noopener noreferrer"&gt;SQL-native standard&lt;/a&gt; for surfacing dbt lineage, metrics, and documentation as plain warehouse tables) is worth tracking. The idea is sound: give every tool that already queries your warehouse access to your metadata without new infrastructure. Atlan's research shows AI agents grounded in governance metadata &lt;a href="https://atlan.com/know/ai-agent/semantic-layer/context-layer-vs-data-catalog-vs-semantic-layer/" rel="noopener noreferrer"&gt;score 38% higher on SQL accuracy&lt;/a&gt; than agents working from schema alone. Whether this standard survives contact with competing vendors is a different question entirely.&lt;/p&gt;

&lt;p&gt;None of this changes the fundamentals of what makes a good data engineering career. Concepts transfer; tools don't. Data modeling, query optimization, understanding why things break. That's the study plan regardless of who owns your transformation layer. If you want to sharpen that foundation, we built snowflake schema practice on datadriven.io specifically because those concepts outlast whatever tool is hot this quarter.&lt;/p&gt;

&lt;p&gt;The merger is done. dbt Core is still open source. The Rust engine is fast. The cost savings are real but metered. The AI features are paywalled. The question that actually matters: 12 months from now, will the feature gap between dbt Core and dbt Cloud be wider or narrower than it is today?&lt;/p&gt;

&lt;p&gt;What's your bet?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>sql</category>
      <category>python</category>
      <category>devops</category>
    </item>
    <item>
      <title>The Junior Data Engineer Job Is Gone. AI Took It.</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Thu, 03 Sep 2026 10:06:32 +0000</pubDate>
      <link>https://dev.to/datadriven/the-junior-data-engineer-job-is-gone-ai-took-it-3g2h</link>
      <guid>https://dev.to/datadriven/the-junior-data-engineer-job-is-gone-ai-took-it-3g2h</guid>
      <description>&lt;p&gt;I've been saying for years that &lt;strong&gt;data engineering&lt;/strong&gt; isn't entry-level. The industry just proved me right in the worst possible way.&lt;/p&gt;

&lt;p&gt;3% of all DE job postings in the US are entry-level. Not 30%. Not 13%. 3%. Out of 6,877 active postings analyzed in May 2026, exactly 219 asked for 2 years of experience or less. That's not a tight market. That's a closed door.&lt;/p&gt;

&lt;p&gt;And the thing that kills me: overall DE &lt;strong&gt;hiring&lt;/strong&gt; is up 23% year-over-year. Companies are posting 1,280 new data engineering positions every week. The field is growing. It's healthy. It's paying well. But every single gain went to mid and senior roles. The &lt;strong&gt;junior data engineer&lt;/strong&gt; job, the one that used to be the on-ramp for every &lt;strong&gt;career&lt;/strong&gt; in this field, is functionally gone.&lt;/p&gt;

&lt;p&gt;AI didn't shrink the junior tier. It deleted it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers Don't Leave Room for Interpretation
&lt;/h2&gt;

&lt;p&gt;Let's break down the current market. 39% of DE postings are mid-level. 31% are senior. 15% are junior. 7% are principal. That 15% sounds survivable until you realize it was roughly double that 2 years ago; &lt;strong&gt;entry level&lt;/strong&gt; tech postings across all fields fell 67% between 2023 and 2024. DE tracked the same curve.&lt;/p&gt;

&lt;p&gt;Stanford's Digital Economy Lab found that junior developer employment for ages 22 to 25 dropped 16% since ChatGPT launched. Workers 30 and older in high-AI-exposure fields saw 6 to 12% &lt;em&gt;growth&lt;/em&gt;. The same technology that's making senior engineers more productive is making junior engineers unemployable.&lt;/p&gt;

&lt;p&gt;Big Tech new-grad hires dropped to 7% of hiring volume. In 2019 it was 30%. The CS Class of 2026 is staring at 6.1% unemployment despite net job growth in the sector. Underemployment for recent college graduates hit 42.5% by Q4 2025.&lt;/p&gt;

&lt;p&gt;54% of engineering leaders plan to hire fewer juniors in 2026. They're not being coy about why: AI copilots let senior engineers cover more ground without backfill.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The field hired 23% more engineers last year and none of them entered at ground level. That's not a hiring dip. That's a structural collapse of the pipeline that creates the next generation.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  AI Owns the Tasks Juniors Used to Learn On
&lt;/h2&gt;

&lt;p&gt;Here's what a junior data engineer used to do in their first year: write staging SQL, scaffold DAGs, build schema mappings, write boilerplate unit tests, handle vanilla ETL from source to warehouse. That was the curriculum. You learned by doing the boring stuff, getting it reviewed, and absorbing context from seniors who'd already made every mistake.&lt;/p&gt;

&lt;p&gt;That curriculum is now a prompt.&lt;/p&gt;

&lt;p&gt;dbt Copilot launched in March 2025. It generates SQL, tests, and documentation for transformation logic. Organizations using AI-powered ETL tools report 40% faster pipeline development and 60% reduction in debugging time. 70% of large enterprise engineering orgs have coding workflows as their highest-penetration LLM use case.&lt;/p&gt;

&lt;p&gt;The work didn't get easier. It got automated. There's a difference.&lt;/p&gt;

&lt;p&gt;When I started, I wrote bad SQL for 6 months before I wrote decent SQL. I wrote staging tables that made no sense. I built DAGs that failed in ways I didn't know were possible. That's how I learned. The feedback loop was: write something wrong, get it reviewed, understand why it's wrong, write it better. Now the machine writes the "decent" version on the first pass. Which sounds great until you realize nobody learned anything.&lt;/p&gt;

&lt;p&gt;70% of hiring managers say AI can do the jobs of interns. 57% trust AI output more than work from recent graduates. I've been on hiring panels where we evaluated candidates, and the uncomfortable truth is: if the only value a junior brings is writing staging SQL and scaffolding DAGs, and a copilot does that in 30 seconds, the business case for that hire evaporates.&lt;/p&gt;

&lt;p&gt;The on-ramp didn't get steeper. The on-ramp got demolished.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Catch-22 That's Killing Early Careers
&lt;/h2&gt;

&lt;p&gt;So here's the loop nobody wants to talk about. You can't get a junior data engineer job because companies aren't posting them. You can't get a senior data engineer job because you don't have experience. You can't get experience because nobody will hire you.&lt;/p&gt;

&lt;p&gt;Time to first job doubled from 4 months in 2022 to 6 to 12 months in 2026. And 48% of visible DE roles are ghost jobs that never actually fill, so the real numbers are worse.&lt;/p&gt;

&lt;p&gt;The traditional path was: get a junior role, spend 2 to 3 years learning systems, get promoted or move to a mid-level role somewhere else, repeat. That path assumed junior roles existed. They don't.&lt;/p&gt;

&lt;p&gt;Meanwhile, bootcamps are still enrolling students for junior data engineer roles the market stopped posting. Bootcamp placement rates sit at 70 to 93%, which sounds great until you learn that "placement" includes analyst roles, customer engineer roles, and data-adjacent positions that aren't DE. The headline number is real. The fine print is doing a lot of heavy lifting.&lt;/p&gt;

&lt;p&gt;I came up through a non-traditional path. No CS degree. Worked my way from analytics through contracting into senior and staff engineering. That path took years and it required companies willing to hire someone without the standard resume. Those companies are harder to find now because AI gave them an alternative to investing in people.&lt;/p&gt;

&lt;p&gt;This is the part that should scare the industry: every junior you don't hire today is a mid-level you won't have in 2028 and a senior you won't have in 2031. The World Economic Forum projects data demand will exceed supply by 30 to 40% by 2027, but the shortage is exclusively for senior talent. We're creating the exact problem we're going to panic about in 3 years.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Gets Callbacks in 2026
&lt;/h2&gt;

&lt;p&gt;I'm not going to sugarcoat this, but I'm also not going to leave you without a path.&lt;/p&gt;

&lt;p&gt;The remaining junior openings don't ask for less skill. They ask for different skill. Python and SQL each appear in 71% of postings; together in 58%. Data pipeline work shows up in 74%. Spark at 38.7%, Snowflake at 29.2%, Databricks at 16.8%. A DE posting in 2026 is a Python job plus a SQL job plus a pipeline job plus a cloud job rolled into one.&lt;/p&gt;

&lt;p&gt;Here's the realistic play: get hired as a data analyst or backend engineer (where entry-level roles still exist at 8% availability, nearly 3 times the DE rate). Do that for 12 to 18 months. Transfer internally to analytics engineer or junior DE. You're doing DE work by month 30. It's slower than the old path. It's also the path that actually works.&lt;/p&gt;

&lt;p&gt;26% of job postings dropped education requirements entirely. That's a real opening for non-traditional backgrounds. But "no degree required" doesn't mean "no skills required." You need a deployed pipeline. Not a tutorial. Not a course certificate. An actual pipeline that fetches data via API, transforms with Python or dbt, loads to a warehouse, and validates schema. Something that runs, breaks, and gets fixed.&lt;/p&gt;

&lt;p&gt;If you're grinding through this market right now, here's my honest advice. Stop studying tools; start studying concepts. Data modeling, query optimization, understanding why things break. These transfer everywhere. The tools change every 18 months. The problems don't. Schema drift, late-arriving data, upstream teams breaking contracts without telling you; these are eternal. If you're prepping for interviews specifically, we put together the data engineer interview questions and answers on datadriven.io for exactly this kind of market, where you can't afford to waste cycles on the wrong prep.&lt;/p&gt;

&lt;p&gt;And stop discounting what you've built. If you've stood up a pipeline, handled data quality issues, worked with stakeholders on requirements, you're not "almost ready." You're doing the job. The title is a formality.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Industry Is Growing. The Ladder Is Broken.
&lt;/h2&gt;

&lt;p&gt;I want to be clear about something: data engineering is not dying. I've been through 3 waves of "data engineering is getting automated away." Still here. Still employed. Still debugging the same categories of problems.&lt;/p&gt;

&lt;p&gt;The field added 23% more roles last year. Senior DE salaries hit $174K median base. The demand-to-supply gap is 3.2 to 1. This is a healthy, expanding profession.&lt;/p&gt;

&lt;p&gt;But a profession that doesn't replenish itself has a shelf life. IBM is trialing the opposite approach, tripling entry-level hiring on the theory that AI-equipped juniors can do formerly senior work. They're redesigning junior roles away from coding toward customer contact and requirements specification. It's an experiment worth watching. Most companies are doing the opposite: squeezing productivity from seniors and hoping the talent pipeline sorts itself out.&lt;/p&gt;

&lt;p&gt;It won't.&lt;/p&gt;

&lt;p&gt;The government noticed (18 months late). The Department of Labor invested $243 million in AI-integrated apprenticeships launching in 2026. Workers with AI competencies earn 56% more than peers without. That's real infrastructure, but it targets 2026 hires after the inflection point already passed. The juniors who needed that ramp in 2024 already chose different fields.&lt;/p&gt;

&lt;p&gt;Junior engineers worry about which tool to learn. Senior engineers worry about which problems to solve. Staff engineers worry about which problems to prevent. Right now, the biggest problem to prevent is an industry that forgot where its seniors come from.&lt;/p&gt;

&lt;p&gt;If you're early in your career and reading this: the path is harder than it was 3 years ago. That's real. But the demand for people who understand data, who can debug pipelines at 2am, who can explain to finance why the numbers don't match; that demand isn't going anywhere. The question is whether the industry builds a new on-ramp before the old generation of seniors starts retiring.&lt;/p&gt;

&lt;p&gt;What's your read? Are companies going to figure this out, or are we going to spend 2028 writing panicked blog posts about the senior DE shortage we manufactured ourselves?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>career</category>
      <category>beginners</category>
      <category>interview</category>
    </item>
    <item>
      <title>Data Engineering Dropped DSA. What Replaced It Is Worse.</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 01 Sep 2026 10:07:33 +0000</pubDate>
      <link>https://dev.to/datadriven/data-engineering-dropped-dsa-what-replaced-it-is-worse-fpo</link>
      <guid>https://dev.to/datadriven/data-engineering-dropped-dsa-what-replaced-it-is-worse-fpo</guid>
      <description>&lt;p&gt;I ran 4 simultaneous data engineering interview loops last year. One company wanted live pair-programming in Cursor with AI explicitly encouraged. The next one sent a 15-hour take-home with a bold-faced warning that any AI usage meant immediate disqualification. The third skipped coding entirely and spent 3 hours on system design. The fourth? Pure SQL, no Python, no design, just "write this query on a whiteboard."&lt;/p&gt;

&lt;p&gt;I prepped for all 4 at the same time. Each required a fundamentally different skill set. I passed 2 and bombed the other 2, and I'm still not sure what the difference was.&lt;/p&gt;

&lt;p&gt;This is what &lt;strong&gt;DSA&lt;/strong&gt; dying actually looks like. Not a clean transition to something better. A collapse into chaos.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Industry Killed DSA and Replaced It With Whatever
&lt;/h2&gt;

&lt;p&gt;The consensus was right: graph traversal and binary search trees were never a meaningful proxy for debugging why a pipeline silently dropped 2M rows last Tuesday. 95% of candidates prefer assessments that mirror actual job scenarios over abstract puzzles. Nobody misses inverting binary trees for a job that's 80% SQL and 20% arguing with upstream teams about schema changes.&lt;/p&gt;

&lt;p&gt;But here's what happened next. Companies dropped DSA and replaced it with whatever their &lt;strong&gt;hiring&lt;/strong&gt; manager felt like that quarter. There was no industry conversation about what comes after. No standard emerged. Instead, we got 4 mutually incompatible formats running simultaneously:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Format 1: Live AI-assisted coding.&lt;/strong&gt; Meta rolled out AI-enabled coding interviews in October 2025. Canva, Rippling, Shopify, and Red Hat followed, redesigning questions to be "more complex, ambiguous, and realistic" to compensate for the AI boost. These rounds test whether you can use tools effectively under pressure. Reasonable in theory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Format 2: AI-banned take-homes.&lt;/strong&gt; 62% of organizations prohibit AI in technical interviews. You get a dataset, a prompt, and 8 to 15 hours of unpaid work. The honor system is the only enforcement mechanism.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Format 3: System design marathons.&lt;/strong&gt; No coding at all. 2 to 4 hours of whiteboarding pipeline architecture, explaining trade-offs, defending decisions. Capital One asks candidates to "design a real-time fraud detection system with 150ms latency" and "process billions of daily transactions." These are Google-caliber questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Format 4: Pure SQL or data modeling.&lt;/strong&gt; Only a third of companies include data modeling rounds, which is wild considering it's the single most transferable skill in the profession. Some shops test nothing but SQL. Others skip it entirely.&lt;/p&gt;

&lt;p&gt;Each format rewards a completely different strength. Take-homes reward AI fluency (whether you admit it or not). Pair-programming rewards narration under stress. System design rewards pattern recall and confidence. Data modeling rewards the thing that actually matters for the job but barely shows up.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;At least DSA was consistent. Now candidates face chaotic, company-specific formats with zero consensus, requiring simultaneous preparation for SQL fluency, data modeling, system design, and Python; covering vastly different skill sets with minimal overlap.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A candidate optimizing for one format will tank another. That's not an assessment of data engineering ability. That's format roulette.&lt;/p&gt;

&lt;h2&gt;
  
  
  The AI Ban That 80% of Candidates Ignore
&lt;/h2&gt;

&lt;p&gt;Here's where it gets truly stupid.&lt;/p&gt;

&lt;p&gt;64% of companies ban AI tools in &lt;strong&gt;interviews&lt;/strong&gt;. Over 50% of candidates use them anyway. Fabric's analysis of 19,368 interviews found 48% of technical candidates showed clear signs of AI assistance. The punchline? 61% of those who cheated still passed.&lt;/p&gt;

&lt;p&gt;Read that again. More than half of the people who broke the rules scored above the approval threshold and received offers.&lt;/p&gt;

&lt;p&gt;This creates a prisoner's dilemma that honest candidates lose. If you follow the rules on a take-home while half the field uses Claude or GPT, you're competing on unequal footing. You're not being evaluated on skill. You're being evaluated on whether you're willing to bend the stated rules when enforcement is nonexistent.&lt;/p&gt;

&lt;p&gt;And the enforcement truly is nonexistent. Interviewing.io ran a blind study: interviewers failed to detect ChatGPT use in 100% of 32 technical interviews. Not "most." All of them. AI detection tools aren't better; one study ran the same essay through 5 different detectors and got scores of 4%, 91%, 12%, 67%, and 38%. An 87-point spread on identical text. That's not detection. That's a random number generator with a corporate logo.&lt;/p&gt;

&lt;p&gt;71% of engineering leaders say AI is making it harder to assess technical skills. Less than 30% have updated their assessment formats or retrained interviewers. So the industry acknowledges the problem, acknowledges it can't detect violations, and continues banning anyway.&lt;/p&gt;

&lt;p&gt;The signal has inverted. The interview no longer measures whether you can build pipelines. It measures whether you'll follow rules that nobody enforces and half the field ignores. Junior candidates cheat at nearly 2x the rate of senior professionals, which makes sense; they have more to lose from an honest showing against AI-polished submissions.&lt;/p&gt;

&lt;p&gt;Companies using verbatim questions see a 73% pass rate. Companies writing custom questions see 25%. That 3x delta tells you exactly what's happening: candidates are reproducing memorized LLM outputs on standard problems. The take-home is dead as a signal mechanism. It just doesn't know it yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAANG Loops at Non-FAANG Salaries
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;interview&lt;/strong&gt; copycat problem has metastasized beyond tech.&lt;/p&gt;

&lt;p&gt;Capital One runs 5-panel hiring loops with system design questions about real-time fraud detection at sub-200ms latency and billion-record daily pipelines. Their Power Day is 2 to 4 hours of back-to-back interviews. The process takes 4 to 8 weeks, sometimes stretching to 10 for engineering roles.&lt;/p&gt;

&lt;p&gt;Google's loop runs 6 to 12 weeks with formal hiring committee review.&lt;/p&gt;

&lt;p&gt;Capital One data engineers earn an average of $133,819. Google pays roughly double for equivalent seniority.&lt;/p&gt;

&lt;p&gt;Same interview. Half the comp. 45% positive candidate experience on Glassdoor, which means more than half of people who go through this gauntlet walk away unhappy. Someone on Blind put it perfectly: "This is on par with Amazon and Google but your compensation package is nowhere near Amazon or Google."&lt;/p&gt;

&lt;p&gt;This pattern repeats across financial services and enterprise tech. Companies copied FAANG's interview structure because it looked rigorous, without copying the compensation that makes candidates willing to endure it. The result: candidate attrition before offers, and the engineers who stick around are the ones with fewer options.&lt;/p&gt;

&lt;p&gt;I've been on both sides of this. I've watched companies run 7-round loops for roles paying $140K, then complain they can't find qualified candidates. You can find them. They're just not willing to do a week of unpaid interviewing for mid-market pay.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Rejections Actually Reveal
&lt;/h2&gt;

&lt;p&gt;A 10-year &lt;strong&gt;data engineering&lt;/strong&gt; veteran with 3 prior FAANG roles passed both SQL and system design at a mid-stage startup. Clean code. Correct surrogate keys. SCD strategy. Idempotency notes. Rejected. The feedback? "Concerns about depth of reasoning."&lt;/p&gt;

&lt;p&gt;This person had built the exact system being asked about. In production. 3 separate times.&lt;/p&gt;

&lt;p&gt;The rejection wasn't about technical ability. It was about narration. The &lt;strong&gt;career&lt;/strong&gt; penalty isn't for not knowing the answer; it's for not performing the answer in the specific way the interviewer expects.&lt;/p&gt;

&lt;p&gt;I failed somewhere around 20 loops before landing multiple offers in a single search. (Yes, I keep count. It's a sickness.) The inflection point wasn't learning new technical material. It was learning to narrate my thinking out loud while solving problems. Meta weights "communication and trade-off articulation" as heavily as technical correctness. That's a learnable skill, but nobody tells you it's the skill being measured.&lt;/p&gt;

&lt;p&gt;The behavioral round has quietly become the round that loses offers. Not the SQL. Not the system design. The "tell me about a time you disagreed with your manager" question that sounds easy until you realize they're evaluating your ability to navigate organizational politics in a 45-minute conversation with a stranger.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Prepare When There Are No Rules
&lt;/h2&gt;

&lt;p&gt;Format research before applying is no longer optional. It's step one.&lt;/p&gt;

&lt;p&gt;Email the recruiter and ask what the loop looks like. Not "what should I prepare?" but "what are the specific rounds, what tools are allowed, and will there be live coding or a take-home?" Read Glassdoor reviews filtered by the hiring manager's name if you can find it. Check LinkedIn for recent hires' backgrounds to reverse-engineer what the team values.&lt;/p&gt;

&lt;p&gt;Hiring timelines now exceed 60 to 90 days with 5 to 7 rounds. You cannot afford to waste 2 months preparing for a system design loop that turns out to be a take-home. 67% of startups explicitly allow AI usage; most enterprises ban it. Know which world you're walking into before you start.&lt;/p&gt;

&lt;p&gt;Then build breadth across the 4 formats:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SQL and data modeling.&lt;/strong&gt; Still the core skill. If you can model a slowly changing dimension and explain why you chose Type 2 over Type 3, you're ahead of most candidates. Data modeling transfers across every tool, every warehouse, every era. The syntax is the easy part; the thinking is what gets tested.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;System design for pipelines, not software.&lt;/strong&gt; Strip back the "design a load balancer" mentality. DEs don't care about reverse proxies. Focus on pipeline architecture: how data flows from source to warehouse, where you'd put quality checks, how you handle late-arriving records, what happens when an upstream team breaks the contract without telling you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Narration under pressure.&lt;/strong&gt; Practice explaining your reasoning out loud while you solve problems. Record yourself. It feels stupid. It works. The engineers failing FAANG loops aren't failing on knowledge; they're failing because they can't communicate while they code. We built the python interview questions for data engineers on datadriven.io specifically because the gap between "I know the concept" and "I can explain it live under pressure" is where most people stall.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Behavioral prep.&lt;/strong&gt; Have 5 stories ready: a time you failed, a time you disagreed, a time you led without authority, a time you made a trade-off under pressure, and a time you debugged something nobody else could find. Structure them. Practice them. This is the round people dismiss and the round that kills offers.&lt;/p&gt;

&lt;p&gt;The data engineering interview process is incoherent right now. That's not changing soon. 26% of job postings don't even mention education requirements. The role itself is still being defined. The tools change every 18 months. The problems don't. Schema drift, late-arriving data, upstream teams breaking contracts without telling you; these are eternal.&lt;/p&gt;

&lt;p&gt;You can wait for the industry to figure out a standard. Or you can treat the chaos as the test it actually is: can you adapt, research, and prepare for ambiguity? Because that's also, coincidentally, the actual job.&lt;/p&gt;

&lt;p&gt;What's the worst interview format you've hit this year, and did the company even tell you what to expect before you showed up?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>interview</category>
      <category>career</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Data Engineers Now Out-Earn Data Scientists. AI Engineers Beat Both.</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Thu, 27 Aug 2026 10:08:48 +0000</pubDate>
      <link>https://dev.to/datadriven/data-engineers-now-out-earn-data-scientists-ai-engineers-beat-both-40fo</link>
      <guid>https://dev.to/datadriven/data-engineers-now-out-earn-data-scientists-ai-engineers-beat-both-40fo</guid>
      <description>&lt;p&gt;3 years ago, a data scientist on my team asked what I made. I told him. He went quiet for about 10 seconds, then said "that can't be right." It was right. I was a senior &lt;strong&gt;data engineer&lt;/strong&gt; clearing more than a senior &lt;strong&gt;data scientist&lt;/strong&gt; with 2 extra years of experience. He'd been told his entire career that the scientist title was the premium path. The market disagreed.&lt;/p&gt;

&lt;p&gt;Here's the thing nobody wants to say out loud: the &lt;strong&gt;data engineer salary&lt;/strong&gt; inversion isn't new. It's been building since 2023. What's new is that it's now undeniable, and a third role, the &lt;strong&gt;AI engineer&lt;/strong&gt;, just showed up and made both of us feel underpaid.&lt;/p&gt;

&lt;h2&gt;
  
  
  Data Engineers Finally Out-Earn Data Scientists. The Numbers Are Clear.
&lt;/h2&gt;

&lt;p&gt;The median base for data engineers in 2026 is $127K. Data scientists? $125K. That's a reversal of everything the industry was told from 2018 to 2020, when data scientists earned 10 to 20% more and every bootcamp on earth was minting "data scientists" like they were going out of style.&lt;/p&gt;

&lt;p&gt;Turns out, they were.&lt;/p&gt;

&lt;p&gt;At the senior level the gap widens fast. Senior data engineers pull $147K to $179K base at mid-market companies, with total comp exceeding $200K at scale. FAANG senior DEs clear $250K to $350K total compensation. Meta data engineers span $168K to $449K depending on level. Databricks engineers average $226K total comp. An analysis of 200K+ job postings shows the median DE posting at $185K, with 75th percentile hitting $221K. Streaming and Spark skills alone add 15 to 25% premiums on top of that.&lt;/p&gt;

&lt;p&gt;The driver isn't some sudden appreciation for pipelines. It's supply and demand; the most boring explanation and the most accurate one. Data science became a bootcamp commodity. The supply of junior data scientists exploded: 8.6% of DS postings are explicitly entry-level, versus 2.6% for &lt;strong&gt;data engineering&lt;/strong&gt;. Meanwhile, infrastructure engineering requires deeper systems knowledge that you can't cram into a 12-week program. The market priced that in.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The salary inversion is real, but the real story isn't the $2K median gap. It's that data engineering has become one of the few technical roles where demand genuinely outpaces supply in 2026.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I've watched data scientists reskilling into data engineering to chase comp bumps. That's the fastest-growing &lt;strong&gt;career&lt;/strong&gt; transition in 2026 according to multiple analyses. 5 years ago that flow went the opposite direction. The industry has fully flipped.&lt;/p&gt;

&lt;h2&gt;
  
  
  The AI Engineer Premium: Real Numbers, Hidden Asterisks
&lt;/h2&gt;

&lt;p&gt;Now here's where it gets interesting. AI engineers earn a median base of $160K, beating data engineers by about $19K. At mid-level, AI engineers pull $140K to $200K total comp. Senior AI engineers? $220K to $310K base, with total comp reaching $340K to $550K. Staff-level packages at places like Databricks hit $480K to $700K.&lt;/p&gt;

&lt;p&gt;And then there are the frontier labs. OpenAI L5 engineers command $1.15M total comp. Anthropic Staff engineers see $1.25M on paper.&lt;/p&gt;

&lt;p&gt;On paper.&lt;/p&gt;

&lt;p&gt;Here's the part that never makes it into the LinkedIn salary flexes: $843K of that Anthropic package is illiquid equity. You can't spend it. You can't sell it. You're betting on an IPO that may or may not happen, at a valuation that may or may not hold. Equity now represents 55 to 70% of frontier lab compensation, up from 35 to 45% in 2024. These are lottery tickets with better odds than Powerball, sure, but they're not paychecks.&lt;/p&gt;

&lt;p&gt;The $200K to $280K total comp range that everyone cites for AI engineers? That's mid-level Big Tech, not frontier. The real premium tier starts at $500K and requires production ML experience that most "AI engineers" don't have. 71% of people hired as AI engineers currently hold titles like "backend engineer" or "infrastructure engineer." The title alone doesn't command the premium. The production chops do.&lt;/p&gt;

&lt;p&gt;One CTO called the AI job title situation in 2026 "the worst naming disaster the industry has produced since we decided DevOps was a person rather than a practice." When a company posts the same job under 3 different titles across 3 req IDs, recruiters overpay the shiniest one by 20 to 40%. That's not a market signal. That's a naming convention doing salary math.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production Skills Pay More Than Any Title
&lt;/h2&gt;

&lt;p&gt;Here's the pattern I keep seeing across every comp dataset: the premium doesn't follow the title. It follows whether you've shipped things to production and kept them running.&lt;/p&gt;

&lt;p&gt;Data engineers now spend 37% of their time on AI projects, up from 19% in 2023. 90% of AI and machine learning projects depend directly on data engineering pipelines. 81% of executives say the data engineer job description has "changed radically due to AI." The role boundaries are blurring so fast that the distinction between "data engineer who works on ML pipelines" and "AI engineer" is mostly a LinkedIn bio decision.&lt;/p&gt;

&lt;p&gt;Production deployment experience commands $15K to $30K additional base salary over pure modeling backgrounds. LLM deployment and fine-tuning expertise adds $20K to $30K. And here's the kicker: the salary premium is nonlinear. Going from zero production skills to one creates the largest compensation bump. Adding a 5th niche skill? Minimal marginal value. This is why generalists with deep production experience often out-earn researchers with exotic modeling expertise but no shipping track record.&lt;/p&gt;

&lt;p&gt;Kafka and Flink production experience alone commands a $15K to $50K premium over the senior band. If you can build and maintain production streaming pipelines, you can basically name your price.&lt;/p&gt;

&lt;p&gt;I've been on hiring panels where a candidate with 3 years of production Kafka got a higher offer than a "senior" generalist with 8 years of experience who'd never operated anything at scale. The market doesn't care about your years. It cares about your reps. The tacit knowledge that comes from systems failing on you in production, from 3am pages and silently dropped records and schema migrations gone sideways, is the one skill set that consistently pays more. It's also the one thing AI can't easily replicate.&lt;/p&gt;

&lt;p&gt;That's the actual differentiator. Not your title. Not your tool list. Whether you've kept something alive in production under real constraints: latency, cost, reliability. A mid-level engineer with 3 years of production streaming experience will out-earn a senior generalist almost every time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should You Rebrand as an AI Engineer? Probably Not.
&lt;/h2&gt;

&lt;p&gt;Here's the career calculus everyone's running right now: do I change my title to "AI Engineer" and ride the wave?&lt;/p&gt;

&lt;p&gt;Short answer: changing your title doesn't change your skills, and hiring managers aren't dumb.&lt;/p&gt;

&lt;p&gt;AI Engineer postings grew 143% year over year. But most successful hires already have non-AI titles. The demand isn't for people called "AI Engineer." It's for engineers who can ship production ML. Rebranding on LinkedIn without rearchitecting your skill mix is resume decoration. You might as well add "thought leader" to your bio while you're at it.&lt;/p&gt;

&lt;p&gt;The smarter move for most data engineers? Go deeper, not wider. Remote data engineering roles now pay $187K median, exceeding San Francisco on-site roles at $179K. Principal IC tracks at top companies pay close to Director money. The DE title isn't a ceiling if you specialize in the things that are actually scarce: streaming infrastructure, ML pipeline ops, data governance at scale (governance managers hit $269K at the 90th percentile, by the way; not exactly a dead-end specialty).&lt;/p&gt;

&lt;p&gt;Data engineers who specialize in DataOps, streaming, or analytics engineering report stronger advancement velocity than engineers who switched titles without upgrading skills. The salary inversion isn't title-driven. It's driven by who controls the critical path to production.&lt;/p&gt;

&lt;p&gt;That said, the window matters. Most analysts think "AI Engineer" as a distinct premium title has a shelf life of 2026 to 2028, maybe 2029. After that, the titles collapse the same way "DevOps Engineer" collapsed into what everyone just calls infrastructure. If you have genuine production ML skills and 4+ years of depth, rebranding now captures the arbitrage. If you're slapping a new title on the same resume, you're wasting everyone's time.&lt;/p&gt;

&lt;p&gt;The real vulnerability isn't being a data engineer instead of an AI engineer. It's being too junior and too generalist. Junior DE postings dropped 67% post-GenAI. The market has stopped onboarding entry-level pipeline builders. Both data engineering and AI engineering are becoming mid-to-senior specialties; the title decision only matters if you've already got 4+ years of depth in one direction.&lt;/p&gt;

&lt;p&gt;Here's what I'd actually do. Pick the specialization that matches your production experience. If you've spent 3 years running Spark jobs and debugging pipeline failures, lean into ML infrastructure. If you've been building streaming systems, that's its own $50K premium without touching the AI engineer title. If you're earlier in your career and wondering where to aim, stack production reps; that's exactly why we built data engineer practice problems on datadriven.io, to give you the scenarios that actually show up in senior interviews and on the job.&lt;/p&gt;

&lt;p&gt;The tools will change. They always do. I've been through 3 waves of "data engineering is getting automated away" and I'm still here, still employed, still debugging the same categories of problems. The concepts transfer. The production instincts transfer. The title on your LinkedIn profile transfers exactly nothing.&lt;/p&gt;

&lt;p&gt;Stop optimizing your job title. Start optimizing the number of production systems you've kept alive at 3am. That's the skill that pays, regardless of what the role is called next quarter.&lt;/p&gt;

&lt;p&gt;Are you seeing the salary inversion play out at your company, or is this still a coastal tech bubble thing? And for anyone who's made the DE-to-AI-engineer jump: was the comp bump real, or did you just trade one set of 3am pages for another?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>career</category>
      <category>interview</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Your Take-Home Was Free Labor. The Data Proves It.</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 25 Aug 2026 10:08:09 +0000</pubDate>
      <link>https://dev.to/datadriven/your-take-home-was-free-labor-the-data-proves-it-pgh</link>
      <guid>https://dev.to/datadriven/your-take-home-was-free-labor-the-data-proves-it-pgh</guid>
      <description>&lt;p&gt;I spent 20 hours on a take-home last year. Built an end-to-end pipeline: ingestion from 3 sources, data modeling with slowly changing dimensions, idempotent loads, test coverage, a README documenting every tradeoff, and a slide deck walking through architecture decisions. The recruiter said it would take "about 4 hours." I knew that was a lie going in; I did it anyway because I wanted the job.&lt;/p&gt;

&lt;p&gt;Got a template rejection 3 weeks later. "We've decided to move forward with other candidates." No feedback. No specifics. No indication that anyone opened the repo.&lt;/p&gt;

&lt;p&gt;Welcome to &lt;strong&gt;data engineering&lt;/strong&gt; hiring in 2026.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 20-Hour Deliverable Nobody Read
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;take-home&lt;/strong&gt; assignment used to be a reasonable signal. Write a script, process some data, show your thinking. 2 hours, maybe 3. That version is dead.&lt;/p&gt;

&lt;p&gt;What replaced it is a full consulting engagement disguised as an assessment. Pipeline design. Multi-source modeling. Testing. Documentation. Edge-case handling. A README with architectural decisions. Sometimes a slide deck. 85% of engineers now encounter take-homes in their &lt;strong&gt;interview&lt;/strong&gt; loops, and the scope has exploded far past what any hiring manager would call "a few hours."&lt;/p&gt;

&lt;p&gt;Industry guidance says cap it at 90 minutes. FAANG managers claim they don't ask candidates to spend more than "a few hours." The reality on the ground? Candidates routinely report 10 to 20 hours of actual effort. One person on Blind documented spending an entire week on a TimescaleDB assignment before receiving a one-line rejection. A DuckDuckGo candidate completed a 7-page project writeup with 15+ requirements, was told the company had a paid policy, and was never compensated.&lt;/p&gt;

&lt;p&gt;Only 4% of engineers report actually getting paid for take-homes. 58% believe they deserve compensation. And 70% complete them anyway because they "really wanted to work at the company." That's not preference. That's coercion by market dynamics.&lt;/p&gt;

&lt;p&gt;Here's the part that should make you angry: only 5.5% of rejected candidates receive useful feedback. 94% want it. 32% were explicitly told their assignment would be discussed at the onsite, then received a template rejection instead. The architecture of no-feedback isn't accidental. It's systematic. Boilerplate rejections prevent learning loops: candidates can't iterate, companies don't field questions, and the next cohort walks into the same opaque process with zero signal from the people who came before them.&lt;/p&gt;

&lt;p&gt;Companies love to cite "legal liability" as the reason they don't give feedback. The actual legal precedent? Zero companies in the US have ever been sued by an engineer who received constructive post-interview feedback. Not one. The legal excuse is an urban myth that persists because it's convenient.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;85% of engineers encounter take-homes. 4% get paid. 5.5% get feedback. The rest donate labor and receive silence.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The &lt;strong&gt;Salary&lt;/strong&gt; Math That Makes It Obscene
&lt;/h2&gt;

&lt;p&gt;Let's run the numbers that nobody at these companies wants you to run.&lt;/p&gt;

&lt;p&gt;Senior data engineers earn $160K to $215K in base &lt;strong&gt;salary&lt;/strong&gt;, with total comp reaching $200K to $300K at strong tech companies. At $250K annually, your hourly rate is roughly $120. A 20-hour take-home means you're donating $2,400 of labor to a company that might ghost you next week. Multiply that across a job search and the math gets ugly fast.&lt;/p&gt;

&lt;p&gt;Companies spend $4,700 per hire on recruitment infrastructure. They'll pay for ATS licenses, recruiter commissions, and job board placements. They'll pay senior engineers to spend 20% of their time reviewing candidates. But compensating the candidate for 20 hours of structured work? That's where the budget apparently runs out.&lt;/p&gt;

&lt;p&gt;A handful of companies now offer $200 to $500 for take-home time. Good for them. But that's also an accidental admission: if you're paying for it, you know it's labor. The 96% of companies that don't pay have made the same calculation and arrived at "we'd rather not acknowledge it."&lt;/p&gt;

&lt;p&gt;And it compounds. Most candidates aren't applying to one company. They're running 5, 10, 15 loops simultaneously. Companies are conducting 42% more interviews per hire than in 2021. Offer rates have collapsed to 38.6%, an 11-year low based on 1.2 million interview reports. At those odds, you're looking at roughly 3 rejections for every offer. If even half those loops include a take-home, a single job search can cost hundreds of hours of unpaid work.&lt;/p&gt;

&lt;p&gt;Under the Fair Labor Standards Act, anyone performing real work that benefits an employer must be paid at least minimum wage, even during a trial. The DOL has already enforced this against companies disguising unpaid labor as "working interviews." Yet no major tech company has been publicly penalized for unpaid engineering trials. The regulatory arbitrage is simple: enforcement costs the candidate more than the company, so nobody challenges it.&lt;/p&gt;

&lt;h2&gt;
  
  
  60 to 90 Days to a Ghost
&lt;/h2&gt;

&lt;p&gt;The take-home is just one piece of the loop. The full picture is worse.&lt;/p&gt;

&lt;p&gt;Data engineering loops now run 5 to 7 rounds. Google's loop stretches 6 to 12 weeks from recruiter call to offer, the longest among FAANG. Industry average across 31 tech roles sits at 6.1 rounds per hire. Engineering roles take 62 days on average, the slowest of any function. And 72% of candidates in active conversation with a hiring team, on a still-open role, have gone 30+ days without any logged recruiter follow-up. The median silence period? 75 days.&lt;/p&gt;

&lt;p&gt;53% of job seekers experienced employer ghosting in 2026, up from 48% in 2025 and 38% in 2024. Three years, straight up. And 48% of formal rejections are bulk job-archive closures where roles get cancelled without individual candidate notification. You didn't fail the loop. The position evaporated mid-process. Nobody told you.&lt;/p&gt;

&lt;p&gt;The timeline trap is vicious: hiring takes 60 to 90 days, but the best candidates exit the market within 10 to 14 days. 42% of candidates abandon the process due to scheduling delays, not failing screens. The bottleneck isn't candidate capability; it's coordination. Companies optimizing for process rigor are losing top talent to competitors who can close in 3 weeks.&lt;/p&gt;

&lt;p&gt;52% of companies openly admit their hiring process is too long. They know. They keep doing it anyway. That's not ignorance; it's organizational lock-in. The hiring committee won't reduce rounds because nobody wants to be the person who approved a bad hire on a shorter loop. So the loop grows, the timeline extends, and the best engineers go somewhere faster while the company's req stays open for another quarter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Senior Engineer, New-Grad Rejection
&lt;/h2&gt;

&lt;p&gt;Here's the story that crystallized the community's anger.&lt;/p&gt;

&lt;p&gt;A senior engineer with 10 years of production experience across 3 FAANG companies submitted a take-home for a mid-stage startup. He built correct surrogate keys, a correct slowly-changing-dimension strategy, documented idempotency, and wrote a detailed tradeoff analysis. When asked about a design decision in the debrief, he answered pragmatically: "I picked customer_id because that's what worked last time I built one of these."&lt;/p&gt;

&lt;p&gt;Rejected. Feedback: "concerns about depth of reasoning."&lt;/p&gt;

&lt;p&gt;He'd built the exact system they were asking about. In production. 3 times. But he didn't articulate a "principled framework." He gave the answer of someone who's done the work, not someone who's rehearsed the vocabulary.&lt;/p&gt;

&lt;p&gt;This isn't an outlier. 28% of job seekers in 2026 cite overqualification as a barrier to employment, equal to those citing underqualification. 12.2% of all rejections explicitly cite "too much experience." At junior-level roles, that number rises to 13%. Companies are screening out experienced engineers not because they lack ability, but because automated systems flag them as flight risks and interviewers penalize pragmatic answers that skip the theoretical preamble.&lt;/p&gt;

&lt;p&gt;A 10-year production engineer reasoning about system-level trade-offs will fail a whiteboard question against a new grad who crammed for 2 weeks. The format doesn't measure what it claims. Interview success correlates with recency of prep, not depth of expertise. I've been on hiring panels where we passed on strong candidates for the dumbest reasons. "Answer was correct but delivery felt rushed." "Used the right approach but couldn't explain why it was right." These are subjective vibes masquerading as evaluation criteria.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should You Walk Away?
&lt;/h2&gt;

&lt;p&gt;59% of candidates walk away when a job listing signals excessive take-home work. That number should be higher.&lt;/p&gt;

&lt;p&gt;Here's my framework. If the take-home exceeds 4 hours of estimated work: ask if they compensate. If they don't, ask yourself whether this company is going to respect your time after they hire you. The answer is usually encoded in how they treat you before they need you.&lt;/p&gt;

&lt;p&gt;If there's no timeline commitment ("we'll get back to you within X days"), that's a signal. If you can't talk to the hiring manager before investing 10+ hours, that's a signal. If the assignment includes deliverables that look suspiciously like a real business problem they're currently trying to solve, that's a signal and possibly a labor law violation.&lt;/p&gt;

&lt;p&gt;The hardest part is that only 6% of candidates refuse take-homes outright. When 94% comply, companies face zero friction. The process persists not because it works, but because unpaid labor is available and nobody's pushing back at scale.&lt;/p&gt;

&lt;p&gt;Your &lt;strong&gt;career&lt;/strong&gt; is a long game. Burning 20 hours on speculative consulting for a company that ghosts you is 20 hours you didn't spend on focused interview prep. Reps on pipeline design, data modeling, and debugging scenarios compound across every loop you enter; that's exactly why datadriven.io has data engineering practice problems targeting what loops actually test instead of what take-homes pretend to. The game is arbitrary. Play the game, win the prize. But don't confuse someone else's consulting project with preparation.&lt;/p&gt;

&lt;p&gt;The process is broken. 52% of companies know it. The data proves it. Nothing changes until candidates stop donating their expertise to a system that values their labor at exactly zero dollars.&lt;/p&gt;

&lt;p&gt;Interviewing is a skill. It's separate from the actual job. Treat prep like a job. But don't treat someone else's job as your unpaid internship.&lt;/p&gt;

&lt;p&gt;What's the worst take-home you've been asked to do, and did you finish it?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>interview</category>
      <category>career</category>
      <category>beginners</category>
    </item>
  </channel>
</rss>
