DEV Community

DataDriven
DataDriven

Posted on

Data Engineering Dropped DSA. What Replaced It Is Worse.

I ran 4 simultaneous data engineering interview loops last year. One company wanted live pair-programming in Cursor with AI explicitly encouraged. The next one sent a 15-hour take-home with a bold-faced warning that any AI usage meant immediate disqualification. The third skipped coding entirely and spent 3 hours on system design. The fourth? Pure SQL, no Python, no design, just "write this query on a whiteboard."

I prepped for all 4 at the same time. Each required a fundamentally different skill set. I passed 2 and bombed the other 2, and I'm still not sure what the difference was.

This is what DSA dying actually looks like. Not a clean transition to something better. A collapse into chaos.

The Industry Killed DSA and Replaced It With Whatever

The consensus was right: graph traversal and binary search trees were never a meaningful proxy for debugging why a pipeline silently dropped 2M rows last Tuesday. 95% of candidates prefer assessments that mirror actual job scenarios over abstract puzzles. Nobody misses inverting binary trees for a job that's 80% SQL and 20% arguing with upstream teams about schema changes.

But here's what happened next. Companies dropped DSA and replaced it with whatever their hiring manager felt like that quarter. There was no industry conversation about what comes after. No standard emerged. Instead, we got 4 mutually incompatible formats running simultaneously:

Format 1: Live AI-assisted coding. Meta rolled out AI-enabled coding interviews in October 2025. Canva, Rippling, Shopify, and Red Hat followed, redesigning questions to be "more complex, ambiguous, and realistic" to compensate for the AI boost. These rounds test whether you can use tools effectively under pressure. Reasonable in theory.

Format 2: AI-banned take-homes. 62% of organizations prohibit AI in technical interviews. You get a dataset, a prompt, and 8 to 15 hours of unpaid work. The honor system is the only enforcement mechanism.

Format 3: System design marathons. No coding at all. 2 to 4 hours of whiteboarding pipeline architecture, explaining trade-offs, defending decisions. Capital One asks candidates to "design a real-time fraud detection system with 150ms latency" and "process billions of daily transactions." These are Google-caliber questions.

Format 4: Pure SQL or data modeling. Only a third of companies include data modeling rounds, which is wild considering it's the single most transferable skill in the profession. Some shops test nothing but SQL. Others skip it entirely.

Each format rewards a completely different strength. Take-homes reward AI fluency (whether you admit it or not). Pair-programming rewards narration under stress. System design rewards pattern recall and confidence. Data modeling rewards the thing that actually matters for the job but barely shows up.

At least DSA was consistent. Now candidates face chaotic, company-specific formats with zero consensus, requiring simultaneous preparation for SQL fluency, data modeling, system design, and Python; covering vastly different skill sets with minimal overlap.

A candidate optimizing for one format will tank another. That's not an assessment of data engineering ability. That's format roulette.

The AI Ban That 80% of Candidates Ignore

Here's where it gets truly stupid.

64% of companies ban AI tools in interviews. Over 50% of candidates use them anyway. Fabric's analysis of 19,368 interviews found 48% of technical candidates showed clear signs of AI assistance. The punchline? 61% of those who cheated still passed.

Read that again. More than half of the people who broke the rules scored above the approval threshold and received offers.

This creates a prisoner's dilemma that honest candidates lose. If you follow the rules on a take-home while half the field uses Claude or GPT, you're competing on unequal footing. You're not being evaluated on skill. You're being evaluated on whether you're willing to bend the stated rules when enforcement is nonexistent.

And the enforcement truly is nonexistent. Interviewing.io ran a blind study: interviewers failed to detect ChatGPT use in 100% of 32 technical interviews. Not "most." All of them. AI detection tools aren't better; one study ran the same essay through 5 different detectors and got scores of 4%, 91%, 12%, 67%, and 38%. An 87-point spread on identical text. That's not detection. That's a random number generator with a corporate logo.

71% of engineering leaders say AI is making it harder to assess technical skills. Less than 30% have updated their assessment formats or retrained interviewers. So the industry acknowledges the problem, acknowledges it can't detect violations, and continues banning anyway.

The signal has inverted. The interview no longer measures whether you can build pipelines. It measures whether you'll follow rules that nobody enforces and half the field ignores. Junior candidates cheat at nearly 2x the rate of senior professionals, which makes sense; they have more to lose from an honest showing against AI-polished submissions.

Companies using verbatim questions see a 73% pass rate. Companies writing custom questions see 25%. That 3x delta tells you exactly what's happening: candidates are reproducing memorized LLM outputs on standard problems. The take-home is dead as a signal mechanism. It just doesn't know it yet.

FAANG Loops at Non-FAANG Salaries

The interview copycat problem has metastasized beyond tech.

Capital One runs 5-panel hiring loops with system design questions about real-time fraud detection at sub-200ms latency and billion-record daily pipelines. Their Power Day is 2 to 4 hours of back-to-back interviews. The process takes 4 to 8 weeks, sometimes stretching to 10 for engineering roles.

Google's loop runs 6 to 12 weeks with formal hiring committee review.

Capital One data engineers earn an average of $133,819. Google pays roughly double for equivalent seniority.

Same interview. Half the comp. 45% positive candidate experience on Glassdoor, which means more than half of people who go through this gauntlet walk away unhappy. Someone on Blind put it perfectly: "This is on par with Amazon and Google but your compensation package is nowhere near Amazon or Google."

This pattern repeats across financial services and enterprise tech. Companies copied FAANG's interview structure because it looked rigorous, without copying the compensation that makes candidates willing to endure it. The result: candidate attrition before offers, and the engineers who stick around are the ones with fewer options.

I've been on both sides of this. I've watched companies run 7-round loops for roles paying $140K, then complain they can't find qualified candidates. You can find them. They're just not willing to do a week of unpaid interviewing for mid-market pay.

What Rejections Actually Reveal

A 10-year data engineering veteran with 3 prior FAANG roles passed both SQL and system design at a mid-stage startup. Clean code. Correct surrogate keys. SCD strategy. Idempotency notes. Rejected. The feedback? "Concerns about depth of reasoning."

This person had built the exact system being asked about. In production. 3 separate times.

The rejection wasn't about technical ability. It was about narration. The career penalty isn't for not knowing the answer; it's for not performing the answer in the specific way the interviewer expects.

I failed somewhere around 20 loops before landing multiple offers in a single search. (Yes, I keep count. It's a sickness.) The inflection point wasn't learning new technical material. It was learning to narrate my thinking out loud while solving problems. Meta weights "communication and trade-off articulation" as heavily as technical correctness. That's a learnable skill, but nobody tells you it's the skill being measured.

The behavioral round has quietly become the round that loses offers. Not the SQL. Not the system design. The "tell me about a time you disagreed with your manager" question that sounds easy until you realize they're evaluating your ability to navigate organizational politics in a 45-minute conversation with a stranger.

How to Prepare When There Are No Rules

Format research before applying is no longer optional. It's step one.

Email the recruiter and ask what the loop looks like. Not "what should I prepare?" but "what are the specific rounds, what tools are allowed, and will there be live coding or a take-home?" Read Glassdoor reviews filtered by the hiring manager's name if you can find it. Check LinkedIn for recent hires' backgrounds to reverse-engineer what the team values.

Hiring timelines now exceed 60 to 90 days with 5 to 7 rounds. You cannot afford to waste 2 months preparing for a system design loop that turns out to be a take-home. 67% of startups explicitly allow AI usage; most enterprises ban it. Know which world you're walking into before you start.

Then build breadth across the 4 formats:

SQL and data modeling. Still the core skill. If you can model a slowly changing dimension and explain why you chose Type 2 over Type 3, you're ahead of most candidates. Data modeling transfers across every tool, every warehouse, every era. The syntax is the easy part; the thinking is what gets tested.

System design for pipelines, not software. Strip back the "design a load balancer" mentality. DEs don't care about reverse proxies. Focus on pipeline architecture: how data flows from source to warehouse, where you'd put quality checks, how you handle late-arriving records, what happens when an upstream team breaks the contract without telling you.

Narration under pressure. Practice explaining your reasoning out loud while you solve problems. Record yourself. It feels stupid. It works. The engineers failing FAANG loops aren't failing on knowledge; they're failing because they can't communicate while they code. We built the python interview questions for data engineers on datadriven.io specifically because the gap between "I know the concept" and "I can explain it live under pressure" is where most people stall.

Behavioral prep. Have 5 stories ready: a time you failed, a time you disagreed, a time you led without authority, a time you made a trade-off under pressure, and a time you debugged something nobody else could find. Structure them. Practice them. This is the round people dismiss and the round that kills offers.

The data engineering interview process is incoherent right now. That's not changing soon. 26% of job postings don't even mention education requirements. The role itself is still being defined. The tools change every 18 months. The problems don't. Schema drift, late-arriving data, upstream teams breaking contracts without telling you; these are eternal.

You can wait for the industry to figure out a standard. Or you can treat the chaos as the test it actually is: can you adapt, research, and prepare for ambiguity? Because that's also, coincidentally, the actual job.

What's the worst interview format you've hit this year, and did the company even tell you what to expect before you showed up?

Top comments (0)