Years ago I inherited a job that had been silently dropping around 40% of its records for 6 months. It didn't crash or alert, and the logs were clean. Someone in finance finally asked why the revenue numbers looked "a little light," and that's how we found out. The fix took an afternoon. Figuring out what had broken took 2 weeks, and the trust took a lot longer to rebuild.
I bring this up because Fivetran just put a price tag on that kind of failure, and it's a big one. Their 2026 enterprise benchmark says pipeline failures expose large companies to roughly $3M a month. In the same stretch of 2026, companies cited AI as the reason for a growing share of tech layoffs. The work that prevents multimillion-dollar outages is the same work getting squeezed, and if you work in data engineering, that tension is also your best career lever right now.
What Fivetran's 2026 benchmark found, and who it surveyed
Fivetran's Enterprise Data Infrastructure Benchmark Report 2026 surveyed 500 senior data and technology leaders in Q4 2025. Every respondent works at an organization with 5,000+ employees, across the US, UK, EMEA, and APAC. The industry mix leans toward financial services (23%), manufacturing (20%), and tech (20%), followed by retail/CPG, healthcare, and a little hospitality. Fivetran reports 95% confidence with a ±4.4% margin of error.
Here are the headline numbers:
- 4.7 pipeline breaks per month on average, rising to 8.3 at the largest enterprises
- 60.4 hours of downtime per month, with each incident taking nearly 13 hours to resolve
- 53% of engineering capacity spent on maintenance instead of new work
- $2.2M per year per enterprise in maintenance labor alone
- An average shop runs 328 pipelines with 35 full-time engineers
- 97% of leaders said pipeline failures have slowed their analytics or AI programs
The math holds together, which I appreciate. 4.7 breaks at roughly 13 hours each comes to roughly 61 hours, which matches the 60.4. Fivetran's press release puts business exposure at $49,600 per hour of downtime. Multiply that by 60.4 hours and you land right around $3M a month. A single incident can reach $1.4M.
Now the part a vendor's marketing team won't put in bold: Fivetran sells managed ELT. The report says legacy and DIY systems break 30% to 47% more often than managed ones, and that managed-platform adopters are "nearly 2x as likely" to beat ROI expectations. That might be true. It's also exactly the finding you'd expect from a company whose business is managed pipelines. Treat the breakage and downtime counts as decent signal and the "buy managed and your problems go away" conclusion as a sales pitch.
I've migrated legacy warehouses at 3am. Managed connectors do remove a whole category of pain, mostly API pagination, auth token rotation, and schema changes in SaaS sources. They don't fix a bad data model, an upstream team renaming a column without telling anyone, or a join that quietly fans out and doubles your revenue.
The number I keep coming back to is the 53%. Of 35 engineers, around 18 spend their week keeping existing pipelines alive. Every senior DE I know would call that number low.
Pipeline failures and layoffs in the same month
Now set that next to the layoff data.
The Challenger, Gray & Christmas May 2026 report has AI cited in about 22% of all US job cuts through May. In May alone it was 40%. Later 2026 coverage of Challenger's numbers puts the year-to-date share around 24% through July, with AI as the most-cited reason for several months in a row. The cumulative counts in secondary write-ups don't fully agree with each other, so I won't pretend to know the exact figure, but the direction is clear. Anyone still quoting "roughly a fifth" is out of date, because the share rose through mid-year.
TechCrunch has kept a running list of large tech layoffs where employers cited AI, including Amazon, Oracle, Microsoft, Meta, and PayPal. Andy Challenger summed it up this way: "Tech remains the center of gravity for this year's cuts, and AI is still the reason companies give."
Keep the phrase "the reason companies give" in mind. What a company says in a press release about a layoff and what actually caused it are often unrelated. I've survived multiple layoff waves. The official reason has always been a story told to the board, and this year the story is AI.
The economics don't work, and I mean that literally. A typical enterprise in Fivetran's sample has $3M a month of downtime exposure, and the people preventing it are a line item in the maintenance budget. Cut 5 of those engineers to save about $1M a year in salary, and you only need a few extra 13-hour incidents to give all of it back. You're also teaching the business that the dashboards can't be trusted, which costs more than any of it.
The labor that keeps $3M a month of downtime contained is the easiest line item to cut and the most expensive one to lose.
I'll be fair to the other side. The same research also shows data engineering postings growing, with demand moving toward people who can handle infrastructure, governance, and cost control. The squeeze lands hardest on generic and lower-tier roles. So the DE role isn't disappearing. It's being repriced, and the people doing the repricing don't fully understand what the job involves.
Where the maintenance hours go
Monte Carlo's 2026 State of Data Quality survey answers most of it. 68% of teams need 4+ hours just to detect an incident, up from 62% in 2022. Average resolution time is 15 hours. And the stat that should embarrass everyone in this industry: 74% of data issues are found by business stakeholders first, up from 47% in 2022.
So for 3 out of 4 issues, the monitoring system is a VP in a Slack channel asking why the numbers look weird. That was my 40% record-drop job exactly, except today it happens at 3x the scale.
That's where the 53% goes. Most of it is detective work: finding which of the 328 pipelines broke, which upstream source changed, which backfill needs to rerun, and which downstream tables are now wrong. The actual fix is usually the shortest part of the incident.
The other surveys in the research fill in the rest:
- dbt's 2026 State of Analytics Engineering report found 72% of data teams prioritize AI coding, while only 24% prioritize AI-assisted pipeline testing and observability. So we're speeding up how fast people write pipelines and barely investing in knowing when they break.
- Atlassian's DX report for Q2 2026 found AI saves engineers 4 to 6 hours a week, yet the innovation ratio barely moved, staying roughly 57% to 58%. The freed-up time gets absorbed by the maintenance backlog.
- Teams with automated observability reportedly resolve incidents about 4x faster. That claim is secondhand, so hold it a bit loosely, but every incident I've worked agrees with the direction.
The tooling is catching up, for what it's worth. Apache Airflow 3.3 shipped AIP-103, a first-class task and asset state store. It persists key-value state across worker crashes and retries, which replaces the XCom hacks that never survived a retry. Apache Iceberg 1.12 replaces position delete files with deletion vectors, compact bitmaps that Dremio says cut read latency by 50% to 80% for high-frequency deletes. That matters a lot for CDC and GDPR workflows.
Both are good releases. Neither changes the fundamentals. Iceberg still needs compaction, snapshot expiration, orphan cleanup, and manifest optimization running in a coordinated way, and if you get orphan retention wrong a mid-write failure can silently corrupt your table. A state store won't help if you don't understand idempotency, because you'll just persist the wrong state faster.
This is the concepts-over-tools argument again. Idempotency, late-arriving data, schema contracts, grain, and backfill strategy work the same in Airflow 2, Airflow 3.3, Dagster, or a cron job someone wrote in 2014. Tools change every 18 months, and the reasons pipelines fail don't.
Your next data engineering interview
Compensation hasn't caught up with reliability work yet. Indeed Hiring Lab's September 2026 analysis found advertised pay in highly AI-exposed occupations up roughly 46% since 2021, compared with 25% for less-exposed fields. The AI premium on identical job titles is about 4.7%, and much larger at senior levels. Mid-level DE base pay sits about $139K median. Meanwhile the people preventing $3M a month in downtime get paid like plumbers. (Plumbers, to be fair, are often doing fine.)
I don't think that lasts, and the hiring signals already show it shifting. Recruiting guides for 2026 now treat "reliability signals" as the minimum: SLAs, monitoring, alerting, backfills, retries, tests, incident response, and postmortems. Senior job postings name observability architecture as a core responsibility, and some staff-level postings run past $260K.
Here's how to use that.
Rewrite your resume around incidents. "Built ETL pipelines using Airflow and Snowflake" is a tool list. "Cut detection time on a revenue pipeline from next-day to 20 minutes by adding row-count and freshness checks" is a story. Don't tell me you "ensured data reliability across a wide array of mission-critical workflows." If I read one more bullet like that I'm putting my fist through drywall.
Have a postmortem ready. Interviewers increasingly ask you to walk through a real incident: how you detected it, what the blast radius was, how you fixed it, and what you changed so it couldn't happen again. Pick your best one and rehearse it until you can tell it in 3 minutes without rambling. If you don't have one, you haven't been doing this long enough, or you were lucky. Either way, go break something in a side project and fix it.
Prep for pipeline architecture. I've watched people with 10 YOE get downleveled because they couldn't explain why their design would survive a late-arriving partition. DEs don't need to whiteboard load balancers. You need to explain how your pipeline handles retries without double-writing, how you'd backfill 90 days without taking down the warehouse, and where your quality checks go. That's the gap we built DataDriven to cover, so for pipeline design practice, try DataDriven: it drills exactly those failure-mode questions, which is where loops separate seniors from everyone else.
Ask about the 53% in your interviews. Ask every hiring manager how many incidents they had last quarter and who found them first. If the answer is "the business usually tells us," you're interviewing for a firefighting job, so price the offer accordingly. If they actually know their numbers, that's a team worth joining.
If you're worried about layoffs, become the person who knows why things break. Every layoff list I've seen spared the engineer who held the context on the scary pipelines. It comes down to what happens when the CFO asks, "who can tell me why the numbers are wrong?" and only one name comes up.
I've been through 3 waves of "data engineering is getting automated away." I'm still here and still debugging the same categories of problems. The Fivetran numbers just put a dollar figure on what every on-call DE already knew: the hard part of the job is keeping things working. The companies cutting that work now will rehire for it at a premium after their first $1.4M incident, and the engineers who can tell a good postmortem story will get those offers.
So here's my question: in your last pipeline incident, who noticed first, your monitoring or someone from the business?
Top comments (0)