<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: DataDriven</title>
    <description>The latest articles on DEV Community by DataDriven (@datadriven).</description>
    <link>https://dev.to/datadriven</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3864671%2F923e8540-fa96-491d-adb6-0e01c42ec26a.png</url>
      <title>DEV Community: DataDriven</title>
      <link>https://dev.to/datadriven</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/datadriven"/>
    <language>en</language>
    <item>
      <title>Data Engineer Salary 2026: Every Survey Is Lying to You</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 21 Jul 2026 17:33:23 +0000</pubDate>
      <link>https://dev.to/datadriven/data-engineer-salary-2026-every-survey-is-lying-to-you-1hkg</link>
      <guid>https://dev.to/datadriven/data-engineer-salary-2026-every-survey-is-lying-to-you-1hkg</guid>
      <description>&lt;p&gt;I pulled &lt;strong&gt;salary&lt;/strong&gt; data from 6 different sources last month. Got 6 different answers. The spread wasn't a rounding error; it was $55,000.&lt;/p&gt;

&lt;p&gt;Career pages from real companies posting real roles showed a median &lt;strong&gt;data engineer&lt;/strong&gt; salary of $185,000. Glassdoor said $134K. ZipRecruiter said $130K. If you're about to walk into a negotiation, which number you believe is the difference between a strong counter and a shrug.&lt;/p&gt;

&lt;p&gt;I've been on both sides of the hiring table at companies whose names you'd recognize. I've watched candidates anchor to the Glassdoor number and leave $40K on the table. I've watched others walk in with career page data, cite it calmly, and get what they asked for, because the hiring manager already knew the budget was there.&lt;/p&gt;

&lt;p&gt;Every major salary survey is structurally wrong. Not "slightly off." Wrong as in they're measuring a different population than the one that's actually getting hired.&lt;/p&gt;

&lt;h2&gt;
  
  
  The $55K Data Engineer Salary Gap Nobody Wants to Explain
&lt;/h2&gt;

&lt;p&gt;An analysis of 244 real job postings pulled from company &lt;strong&gt;career&lt;/strong&gt; pages in 2026 shows a median data engineer salary of $185,000. Remote roles median even higher at $187,000. San Francisco, weirdly, comes in at $179,000.&lt;/p&gt;

&lt;p&gt;Now compare that to the survey platforms. ZipRecruiter: $129,716. Glassdoor: $133,861. Indeed: $136,776. Each one sits $50K+ below what companies are actually posting when they're trying to fill seats.&lt;/p&gt;

&lt;p&gt;This is not a disagreement about methodology. It's a $55K chasm caused by fundamentally different measurements. Career pages capture what a company will pay &lt;em&gt;right now&lt;/em&gt; to hire someone. Surveys capture what a mixed bag of respondents reported earning at some point in the recent past. These are different questions producing different answers, and most people don't realize they're looking at the wrong one.&lt;/p&gt;

&lt;p&gt;The gap gets worse when you factor in that only 14% of tech job postings even disclose salary. When companies &lt;em&gt;do&lt;/em&gt; post numbers, they tend to be the ones with competitive budgets. The thousands of postings with no salary listed? Those are the ones dragging survey medians down through omission.&lt;/p&gt;

&lt;p&gt;And then there's FAANG, which breaks every survey completely. Meta data engineers: $322K median total comp. Google: $276K. Netflix: $565K in straight cash. An E5 at Meta (senior level) pulls $229K base + $222K annual stock vest + $27K bonus. That's $478K total. Netflix L5 clears $550K with no equity complexity at all. None of these numbers exist in Glassdoor. They're invisible to traditional survey methodology because the people earning them don't fill out surveys.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Every Data Engineering Salary Survey Gets It Wrong
&lt;/h2&gt;

&lt;p&gt;Here's the dirty secret about salary surveys: they don't measure the market. They measure &lt;em&gt;whoever decided to fill out a survey&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;And who fills out salary surveys? Not the L5 at Netflix clearing $550K in cash. Those people have zero incentive to spend 15 minutes on a Glassdoor form. The people who fill out surveys are disproportionately earlier in their &lt;strong&gt;career&lt;/strong&gt;, disproportionately frustrated with their pay (59% of tech workers report feeling underpaid), and disproportionately concentrated in a handful of metros.&lt;/p&gt;

&lt;p&gt;This creates 4 compounding biases that make every number you see structurally wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Self-selection bias.&lt;/strong&gt; Crowdsourced salary data is driven by who shows up. Workers who feel underpaid submit to complain. Workers who feel overpaid submit to boast. The actual median; the people in the middle? They're doing their jobs. One audit found 43% of crowdsourced salary submissions were off by 15%+ from market benchmarks. That's not noise. That's a broken instrument.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Geographic skew.&lt;/strong&gt; San Francisco Bay Area tech salaries run 126.6% of the national average. SFBA software engineers earn a median $233K while national surveys report $125K to $135K. Every survey oversamples coastal hubs because that's where the respondents are, which pulls the number in directions that don't represent the national market or the remote market where the money increasingly lives.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Title dilution.&lt;/strong&gt; "Data engineer" in 2026 could be a warehouse analyst at $90K, an ML platform engineer at $200K+, or a senior data architect at $250K+. Surveys that report a single median for this title are averaging incomparable roles. It's like reporting the "average vehicle price" across sedans, dump trucks, and Ferraris.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The equity invisibility tax.&lt;/strong&gt; Surveys almost never disaggregate base from total comp. A data engineer at FAANG might show $230K base, but the $170K in annual equity vesting and $25K bonus bring real compensation to $425K. Meanwhile, a survey respondent at a non-tech enterprise sees base ≈ total comp. Mashing these together into one "average" is statistical malpractice.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The survey is measuring what people accepted 2 years ago. The career page is measuring what companies will pay today. If you're negotiating tomorrow, only one of those numbers is useful.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here's the thing about Glassdoor specifically: it wraps self-reported data in additional comp estimates, which inflates some figures while dragging others down. PayScale skews early-career because of who fills out their forms. ZipRecruiter reflects what's &lt;em&gt;posted&lt;/em&gt;, not what's &lt;em&gt;accepted&lt;/em&gt;. Each source surveys a different crowd and measures a different thing. Workers making $250K+ rarely respond to public salary surveys at all; privacy risk, employer visibility, lack of motivation. The top 15% of earners are systematically underrepresented, pulling reported medians down 10% to 15% compared to actual compensation.&lt;/p&gt;

&lt;p&gt;There is no unbiased source. But some sources are less wrong than others.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Entry-Level Collapse That Broke Every Average
&lt;/h2&gt;

&lt;p&gt;Only 3% of data engineer postings in 2026 are entry-level. Down from 10% to 15% historically. Junior &lt;strong&gt;data engineering&lt;/strong&gt; postings fell 67% post-GenAI, with entry-level hiring collapsing 73% year-over-year between late 2023 and late 2025.&lt;/p&gt;

&lt;p&gt;66% of CEOs are actively freezing entry-level headcount. Not pausing. Freezing. They're trading junior hires for "judgment hires": mid and senior engineers who can architect systems, debug production failures, and make governance decisions that AI can't touch. AI automated the boilerplate: staging SQL, scaffolded DAGs, schema mappings. It did not automate architecture, governance, debugging judgment, or cost optimization.&lt;/p&gt;

&lt;p&gt;The total data engineer market still grew 23% year-over-year in headcount. The global DE services market hit $105 billion growing at 15% CAGR. But all of that growth is seniority-weighted. Companies are hiring &lt;em&gt;more&lt;/em&gt; data engineers; they're just not hiring juniors.&lt;/p&gt;

&lt;p&gt;This does 2 things to salary data. First, it pulls every average up mechanically. When the bottom 10% to 15% of earners vanishes from the hiring pool, the median jumps without anyone getting a raise. The $185K career page median isn't inflated; it's accurate for the population that's actually getting hired. The survey-reported $130K is understated because it's sampling a truncated pool that includes people who accepted junior rates 3 years ago.&lt;/p&gt;

&lt;p&gt;Second, it creates a bifurcated market that a single median can't capture. Junior data engineer ranges sit at $72K to $97K on ZipRecruiter. But Glassdoor reports $126K average for the same title; a 75% variance driven by geographic and sample bias. Base salaries fell 15% to 25% below 2022 peaks, but this hit juniors disproportionately. AI/ML specialists command 30% to 50% premiums over generalists. If you can run production Kafka pipelines, you're in a fundamentally different market than someone looking for their first role.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Companies are advertising junior roles, then quietly filling them with experienced engineers. This isn't a hiring freeze; it's a bait-and-switch.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The uncomfortable math: if companies stop hiring juniors today, they're engineering a senior shortage in 5 to 10 years. But CFOs don't optimize for 2031. They optimize for this quarter's headcount target.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to Actually Bring to Your Next Interview and Negotiation
&lt;/h2&gt;

&lt;p&gt;60% to 70% of candidates accept the first offer without countering. Those who negotiate with market data see 15% to 20% increases on average; about $24K median increase in tech.&lt;/p&gt;

&lt;p&gt;The difference between a good &lt;strong&gt;salary&lt;/strong&gt; negotiation and a bad one is which data you walk in with.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Career page postings&lt;/strong&gt; for your target company and comparable companies. These reflect current budgets, not historical averages. If the posting shows $170K to $210K, your anchor is the 75th percentile, not the midpoint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Levels.fyi&lt;/strong&gt; for tech-specific roles, especially FAANG. The data disaggregates base, equity, and bonus, which matters when total comp runs 2x to 3x base. And here's what most people miss about equity: it's renewing, not depreciating. An L5 engineer's compensation contains overlapping tranches. Initial grant (years 1 to 4), year-2 refresher (years 2 to 5), year-3 refresher (years 3 to 6). This creates a $125K to $150K annual equity floor &lt;em&gt;after&lt;/em&gt; the initial grant vests. Ask explicitly: "What's the typical annual refresher equity grant?" That's the question that separates people who understand comp from people who don't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Robert Half and Motion Recruitment&lt;/strong&gt; salary guides, which segment by level and geography. Robert Half reports $127K to $180K entry-level, $160K to $215K senior. Tighter ranges, more useful than a single median.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Glassdoor and ZipRecruiter&lt;/strong&gt; as a floor, not a ceiling. If the survey says $130K, treat that as the minimum. Your job is to prove you're worth the career page number.&lt;/p&gt;

&lt;p&gt;Skill premiums stack and they're documented. Kafka/streaming production experience: $15K to $50K over senior baseline. AWS Data Analytics Specialty certification: +$18K. Spark expertise: 15% to 25% premium. Both batch and streaming? 30% to 50% salary jump. Streaming carries the steepest premium because production Kafka is hard to learn without a production Kafka environment; the barrier to entry is structural, not educational.&lt;/p&gt;

&lt;p&gt;Don't forget the base salary multiplier: your bonus is usually a percentage of base. The higher you negotiate base, the higher every downstream calculation. One negotiation compounds for years.&lt;/p&gt;

&lt;p&gt;And here's the part nobody tells you: &lt;strong&gt;interview&lt;/strong&gt; prep and negotiation prep are the same skill. The better you perform in the loop, the stronger your leverage on the offer. When I was grinding through 20+ loops in a single job search, the difference between the lowball offers and the strong ones tracked almost perfectly with how well I'd prepared for each company's process. That's the problem we set out to solve with &lt;a href="https://datadriven.io" rel="noopener noreferrer"&gt;DataDriven&lt;/a&gt;; when someone says i used DataDriven for data pipeline interview questions, those reps covered the patterns that actually show up in loops, not generic textbook exercises.&lt;/p&gt;

&lt;p&gt;Colorado's pay transparency law alone pushed posted salaries up 3.6%. As more states mandate disclosure, the gap between survey data and reality will shrink. But right now, in mid-2026, the data engineering salary market has a $55K information asymmetry. The side you're on determines whether you negotiate from strength or from a number that was wrong before you opened your mouth.&lt;/p&gt;

&lt;p&gt;The tools change. The surveys will keep being wrong in the same ways for the same reasons. Learn which numbers to trust, walk in with the right data, and stop letting a Glassdoor screenshot be the reason you leave 5 figures on the table.&lt;/p&gt;

&lt;p&gt;What's the biggest gap you've seen between what a survey reported and what you actually got offered?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>career</category>
      <category>interview</category>
      <category>beginners</category>
    </item>
    <item>
      <title>The AI Cheat Tool Your Interview Cannot See</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Thu, 16 Jul 2026 10:12:37 +0000</pubDate>
      <link>https://dev.to/datadriven/the-ai-cheat-tool-your-interview-cannot-see-jkc</link>
      <guid>https://dev.to/datadriven/the-ai-cheat-tool-your-interview-cannot-see-jkc</guid>
      <description>&lt;p&gt;I've been on &lt;strong&gt;hiring&lt;/strong&gt; panels where we spent 45 minutes convinced a candidate was sharp. Articulate answers. Clean code. Solid reasoning. Then in the debrief, someone pulled up the recording and timed the responses. Every answer: 4 seconds. Easy question, hard question, curveball follow-up. 4 seconds flat. Humans don't think like that. Humans stumble on hard problems and breeze through easy ones. This person's cadence was perfectly uniform. I'd love to tell you I spotted it in real time. I didn't. Nobody on the panel did.&lt;/p&gt;

&lt;p&gt;That was 8 months ago. The tools have gotten significantly better since.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Overlay That Broke Data Engineering Hiring
&lt;/h2&gt;

&lt;p&gt;There's a new category of &lt;strong&gt;AI cheating&lt;/strong&gt; tools that most interviewers don't know exist. Cluely, Interview Coder, Final Round AI; these aren't browser tabs a candidate Alt-Tabs to when you're not looking. They're &lt;strong&gt;invisible GPU overlays&lt;/strong&gt; that use DirectX (Windows) and Metal (macOS) hooks to render LLM-generated answers directly on the candidate's monitor, beneath the capture layer that Zoom, Teams, and Google Meet use for screen sharing.&lt;/p&gt;

&lt;p&gt;In plain English: when you screen-share on Zoom, the application captures pixels from a specific layer of the OS graphics pipeline. These overlay tools render their content below that layer, in the GPU's local frame buffer. The interviewer sees a clean IDE. The candidate sees the IDE plus a floating panel with AI-generated code and explanations. Those pixels literally don't exist in the video stream that gets encoded and transmitted.&lt;/p&gt;

&lt;p&gt;This isn't a proof of concept from a security researcher. Cluely pulled 70,000 signups in its first week. It's a consumer product with standard SaaS pricing: free tier at 5 responses per day, $20/month for Pro, $75/month for the "undetectability" tier that does the GPU-level rendering. Cheating on your &lt;strong&gt;data engineering&lt;/strong&gt; interview now costs less than a monthly gym membership.&lt;/p&gt;

&lt;p&gt;The overlays don't trigger tab-switch alerts. They don't appear in keystroke logs. They don't show up in screen recordings. The proctoring vendors claiming 85-95% detection effectiveness? They're scanning for browser tab switches and copy-paste events. They're watching the wrong layer entirely. 6 new overlay tools emerged in 2025 alone, plus at least 3 open-source clones.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Screen-share monitoring is security theater. The pixels the interviewer sees are not the pixels the candidate sees.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Half Your Tech Candidates Are Cheating. Most of Them Pass.
&lt;/h2&gt;

&lt;p&gt;Fabric analyzed 19,368 &lt;strong&gt;interviews&lt;/strong&gt; between July 2025 and January 2026 across 50+ companies. The numbers are bleak.&lt;/p&gt;

&lt;p&gt;38.5% of all candidates triggered cheating flags. For technical roles, that number hit 48%. Sales roles? 12%. The roles where the stakes are highest and the questions are most googlable are the roles getting gamed hardest. This inverts the assumption that technical interviews are somehow more honest by nature. They're not. They're just more automatable.&lt;/p&gt;

&lt;p&gt;The acceleration is what gets me. Cheating went from 9% in July 2025 to 45% by September. 3x in 3 months. Then it plateaued through January; not because people stopped, but because it saturated. When nearly half your candidate pool is using AI assistance, you don't have a cheating problem. You have a broken measurement system.&lt;/p&gt;

&lt;p&gt;The part that should make every &lt;strong&gt;hiring&lt;/strong&gt; manager stop and think: 61% of flagged cheaters scored above the passing threshold and advanced in the pipeline. These aren't marginal candidates scraping by. They're clearing the bar comfortably, because the AI is genuinely good at answering the questions we ask. 61% is not a detection gap. It's an action gap. Companies flag candidates and hire them anyway because the score looked fine.&lt;/p&gt;

&lt;p&gt;Junior candidates (0 to 5 years of experience) cheat at nearly double the rate of seniors. The people with the least context to evaluate whether an AI-generated answer is even correct are the ones relying on it most. It's a desperation play: they know AI fills knowledge gaps faster than grinding prep, and they're probably right. The incentive structure is broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  Banning AI Doesn't Fix the Architecture
&lt;/h2&gt;

&lt;p&gt;64% of companies ban AI in interviews. 80% of candidates use it anyway on take-homes. That's not a policy failure; that's a policy being laughed at.&lt;/p&gt;

&lt;p&gt;The detection layer makes it worse. AI detection tools produce essentially random results: the same essay scored 4%, 91%, 12%, 67%, and 38% across 5 different detectors. OpenAI killed its own classifier in July 2023 because it was only 26% accurate. Stanford found a 61.2% false-positive rate on essays from non-native English speakers versus 5.1% for native speakers. You're not catching cheaters. You're penalizing people who learned English as a second language.&lt;/p&gt;

&lt;p&gt;The industry response has been a stampede back to in-person interviews. They jumped from 5% to 30% in a single year. Google and McKinsey mandated in-person rounds in mid-2025, explicitly citing AI fraud. 72% of recruiting leaders say fraud prevention is the driver.&lt;/p&gt;

&lt;p&gt;The problem: in-person doesn't scale. It's expensive, exclusionary, and locks out remote candidates, international candidates, and anyone who can't fly to your office for a day. The least scalable option is the one everyone's reaching for. That's not a fix; it's a retreat to 2019.&lt;/p&gt;

&lt;p&gt;Take-homes are dead in their current form. Anthropic's own engineering team found that Claude Opus "matched top candidates" on take-home assessments and there was "no longer a way to distinguish between the output of top candidates and the most capable model." When the company building the AI tells you their model passes your take-home, believe them.&lt;/p&gt;

&lt;p&gt;Eye tracking? Gaze monitoring? Also crumbling. NVIDIA Maxine synthesizes natural gaze. Candidates keep their eyes near the camera while reading overlays at the screen periphery. HireVue discontinued facial analysis in 2021 after finding facial expression data contributed less than 0.25% to performance predictions. The biosurveillance approach is simultaneously invasive and useless.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You can ban AI in interviews. You can't ban the architecture it runs on. The enforcement tools are more invasive than the cheating, and they still don't work.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Interview That Survives
&lt;/h2&gt;

&lt;p&gt;The strongest signal isn't software. It's one follow-up question.&lt;/p&gt;

&lt;p&gt;AI generates polished, multi-layered code. But when you ask a candidate to explain line 7, or walk through why they chose that join strategy, or refactor their solution with a new constraint, the gap between "read this off a screen" and "actually understand it" becomes obvious in seconds. If you can't explain line 7, you didn't write line 7.&lt;/p&gt;

&lt;p&gt;Consistent response timing is the second strongest behavioral fingerprint. Humans pause longer on hard questions. They stammer, correct themselves, say "let me think" with variable duration. AI overlay tools show uniform 3 to 5 second latency regardless of difficulty, because the pipeline is always: audio capture, transcription, LLM inference, render. That flatness is detectable. But only about 30% of interviewers actively monitor response timing. The other 70% miss it entirely.&lt;/p&gt;

&lt;p&gt;System design remains the most AI-resistant format because it's fundamentally discursive. You can't read a system design answer off an overlay, because system design isn't a question with a fixed answer. It's a conversation that shifts when the interviewer changes constraints, pushes back on choices, and asks "what happens when this fails?" No overlay tool fakes that in real time. Not yet.&lt;/p&gt;

&lt;p&gt;For &lt;strong&gt;data engineering&lt;/strong&gt; interviews specifically, the path forward is clearer than for most roles. Pipeline architecture, data modeling exercises, debugging scenarios. "Here's a pipeline that silently dropped 2M rows last Tuesday. Walk me through how you'd find the problem." That question has no clean overlay answer because the solution depends on follow-ups about the specific system, the specific data, the specific failure mode. The actual job is debugging, not building; might as well test for it.&lt;/p&gt;

&lt;p&gt;This is also why concepts matter more than tools, and always have. If your &lt;strong&gt;interview&lt;/strong&gt; tests whether someone can write a Spark transformation, an overlay solves it in 4 seconds. If it tests whether someone understands why a pipeline breaks when upstream schema changes violate a downstream join contract, you're testing something AI can't fake. Concepts transfer across tools; syntax doesn't, and that gap is exactly why we built databricks interview prep with datadriven around architecture walkthroughs and data modeling, not the kind of timed problems an overlay solves in 4 seconds.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;career&lt;/strong&gt; implications cut both directions. If you're a candidate who genuinely knows your stuff, push for system design and pair-programming rounds. Those formats favor you. If a company's entire loop is timed LeetCode and async take-homes, their signal is already compromised, and your real expertise is competing against someone paying $75/month for invisible help.&lt;/p&gt;

&lt;p&gt;The companies that adapt their process will hire better. The ones clinging to LeetCode mediums and take-homes will keep onboarding engineers who can't debug a broken DAG in their first week, then blame the candidate instead of the process.&lt;/p&gt;

&lt;p&gt;What's the most egregious cheating you've seen in an interview, on either side of the table?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>interview</category>
      <category>career</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Half of Data Engineering Jobs on LinkedIn Aren't Real</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 14 Jul 2026 10:06:48 +0000</pubDate>
      <link>https://dev.to/datadriven/half-of-data-engineering-jobs-on-linkedin-arent-real-3gfd</link>
      <guid>https://dev.to/datadriven/half-of-data-engineering-jobs-on-linkedin-arent-real-3gfd</guid>
      <description>&lt;p&gt;I applied to 47 data engineering jobs in a three-week stretch last year. Heard back from nine. Got interviews at four. Two of those companies had already filled the role before my first call. One admitted the headcount was "paused indefinitely." The fourth gave me a verbal offer that evaporated when the hiring manager left. That's the &lt;strong&gt;job market&lt;/strong&gt; in 2026. You're not failing; you're playing a rigged game.&lt;/p&gt;

&lt;p&gt;Here's the part that should make you angry: companies are publicly claiming &lt;strong&gt;data engineering&lt;/strong&gt; hiring is up 23% year-over-year. That number is real. It's also one of the most misleading statistics in tech right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Paradox: Hiring Is Up, But the Ladder Is Gone
&lt;/h2&gt;

&lt;p&gt;Data engineering hiring grew 23% YoY. &lt;strong&gt;Entry level&lt;/strong&gt; data engineering roles fell 67% since AI went mainstream. Both of these things are true at the same time.&lt;/p&gt;

&lt;p&gt;The growth is exclusively senior hires. Companies aren't expanding their data teams; they're replacing junior pipelines with senior architects who can ship AI-ready infrastructure on day one. Only 2.3% of DE job postings target entry-level candidates with under two years of experience. The most common requirement? Four to six years, appearing in 11% of postings.&lt;/p&gt;

&lt;p&gt;This isn't a downturn. It's a reclassification. What used to be "entry-level" got relabeled as "mid-level with 4+ years." The ladder didn't break; it got pulled up.&lt;/p&gt;

&lt;p&gt;Junior developer postings across all of tech collapsed 60% between 2022 and 2024. Data engineering followed the same trajectory, just quieter. And here's the twist nobody talks about: job postings labeled "entry-level software engineer" grew 47% between October 2023 and November 2024, but actual hiring into those levels dropped 73% in the same window. Companies are advertising junior roles and filling them with experienced engineers. The title is a lie.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The market didn't grow 23%. It compressed vertically. The number went up; the ladder got removed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Meanwhile, 66% of CEOs surveyed are freezing or cutting hiring through the rest of 2026 while simultaneously betting billions on AI infrastructure. LLM engineering skills in DE job postings spiked 300% in a single quarter, from 3% to 12%. The role is being rewritten in real time, and the rewrite doesn't include a chapter for people just starting out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Half the Job Market Is a Mirage
&lt;/h2&gt;

&lt;p&gt;Let me say this plainly: &lt;strong&gt;ghost jobs&lt;/strong&gt; now account for 48% of tech job postings. Nearly half the roles you see on LinkedIn aren't real. They're not stale listings that HR forgot to take down (though some are). They're deliberate.&lt;/p&gt;

&lt;p&gt;40% of tech companies posted fake jobs in the past year, and 79% of those listings were still active at the time of the survey. This isn't accidental. It's infrastructure.&lt;/p&gt;

&lt;p&gt;Here's why companies do it. 62% of &lt;strong&gt;hiring&lt;/strong&gt; managers admitted they post ghost jobs to make current employees feel replaceable. 43% cited signaling company growth to investors and board members. Nearly 60% collected resumes with no intention to hire immediately; they call it "talent pool management." The idea originates from HR (37%), senior management (29%), or executives (25%). Hiring managers, the people who actually need to fill roles, typically don't originate it.&lt;/p&gt;

&lt;p&gt;93% of HR professionals engage in posting ghost jobs: 45% regularly, 48% occasionally. And 96% of recruiters use automated software to repost listings on a schedule, so jobs disappear and reappear with identical descriptions, the clock resetting every 30 to 90 days. The ATS never stops accepting applications even after the requisition is dead. If you got an automated rejection two to four hours after applying, that's not a keyword mismatch; that's a closed role running on autopilot.&lt;/p&gt;

&lt;p&gt;The financial damage is real. 72% of job seekers report mental health damage from the application process. 37% suffer direct financial losses averaging $500 to $2,500. And the 47% of tech professionals actively job-hunting in 2026 (up from 29% last year) means there's an endless supply of desperate applicants feeding the ghost job ecosystem. Companies have no incentive to stop.&lt;/p&gt;

&lt;p&gt;The worst offenders aren't FAANG and they aren't tiny startups. Companies with 1,001 to 5,000 employees post ghost jobs at nearly a 25% rate, the highest of any company size. That's the Series C through E band where CFOs tighten headcount while board pressure demands growth signals. If you're an early-career engineer, that's the exact cohort you're probably targeting. You picked the most deceptive segment of the market.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Real Postings Actually Want
&lt;/h2&gt;

&lt;p&gt;So what do the legitimate data engineering roles look like in 2026? Different. Fundamentally different from two years ago.&lt;/p&gt;

&lt;p&gt;Python (70%) and SQL (69%) are still non-negotiable. That hasn't changed. Everything else has. Machine learning appears in 29.9% of postings. Kafka shows up in 24%. CI/CD in one out of six. Apache Spark still dominates at 38.7%, but Snowflake (29.2%) and Databricks (16.8%) are carving out separate tiers. Companies don't want generalists who can learn their stack; they want people already fluent in it.&lt;/p&gt;

&lt;p&gt;The role has absorbed platform engineering, DevOps integration, ML pipeline support, and governance orchestration into a single position. Data engineers must prepare data for AI use cases, collaborate with ML engineers, and understand feature stores, experimentation, and model serving. That sentence would have been nonsensical in a DE job description 18 months ago. Now it's baseline.&lt;/p&gt;

&lt;p&gt;26% of job postings don't mention education requirements at all. That sounds like a window for self-taught engineers, and it is, but it's being filled by mid-career pivots with adjacent experience, not by people fresh out of bootcamp. The real barrier isn't credentials; it's that bootcamp curricula teach isolated SQL and Python while jobs demand LLM-aware pipelines, regulatory audit trails, and 99.95% uptime infrastructure. The skill tree forked, and the entry-level branch got pruned.&lt;/p&gt;

&lt;p&gt;The median salary sits at $131K to $135K, which sounds great until you realize it's skewed by the shift toward senior talent. Senior contract data engineers command $150 to $185 an hour. Specialized AI-infrastructure architects bill $220 to $400. The floor rose because the people standing on it changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the Actual Jobs Are (and How to Stop Wasting Time)
&lt;/h2&gt;

&lt;p&gt;Real jobs exist. They're just not where most people are looking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contract and platform migration work&lt;/strong&gt; is where the volume is. Over 3,300 data migration jobs and 6,600+ technical data migration roles are active right now, mostly 60 to 90 day engagements for companies rebuilding data stacks for cloud migration and AI readiness. Snowflake and dbt expertise commands premium contract rates. These aren't glamorous FTE positions with equity; they're sprint work. But they're real, they pay, and they build the exact résumé signals that get you into full-time senior roles later.&lt;/p&gt;

&lt;p&gt;Databricks alone has 840+ open roles. But here's the catch: that's the vendor hiring, not the customers. If customers were expanding data teams, they'd be hiring people to &lt;em&gt;use&lt;/em&gt; Databricks. Instead, Databricks is hiring its own engineers to do POC work for under-resourced customers. Tool adoption isn't translating to team growth at the companies actually using the tools.&lt;/p&gt;

&lt;p&gt;To spot a real posting versus a ghost, here's what I actually look at:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check the company's own careers page.&lt;/strong&gt; Two minutes on their site is still the best ghost job filter available. If the role isn't listed there, it's phantom. LinkedIn aggregates are noise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Look at the posting age.&lt;/strong&gt; Tech roles typically fill in 30 to 45 days. If a listing has been open for 90+ days, something is wrong. Jobs posted 30+ days show a 30% chance of never resulting in a hire.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Salary range omission is a red flag.&lt;/strong&gt; 16+ states and D.C. now mandate pay disclosure. Listings that dodge it in those jurisdictions are either non-compliant or not real. 44% of candidates won't apply without a range; legitimate employers know this.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Watch for reposting patterns.&lt;/strong&gt; Same title, same description, fresh date. That's the ATS auto-renewing a dead requisition. The clock reset doesn't mean new hiring intent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Named hiring manager &amp;gt; generic HR.&lt;/strong&gt; If the posting names a specific manager or team lead, someone is actually waiting for a hire. Generic "talent team" postings with no team context are pipeline collection.&lt;/p&gt;

&lt;p&gt;The honest advice for early-career engineers: stop applying to 50 listings a week and start being surgical. Target companies under 1,000 or over 10,000 employees (lower ghost rates). Prioritize contract work that builds architectural experience. Get reps on the stuff that actually separates candidates in interviews: data modeling, pipeline architecture, system design thinking. That's exactly why we built &lt;a href="https://datadriven.io" rel="noopener noreferrer"&gt;datadriven.io&lt;/a&gt;; DataDriven is good for data engineer interview questions that test the concepts behind the tools, not trivia about Spark APIs nobody remembers anyway.&lt;/p&gt;

&lt;p&gt;The entry-level data engineering path isn't dead. It's been rerouted. The reliable path now is analyst or backend engineer first, then internal transfer. That sounds frustrating, and it is. But it's also how most of us got here. I didn't start in data engineering. I started outside of tech entirely. The path was never a straight line; we just pretended it was for a few years when hiring was hot.&lt;/p&gt;

&lt;p&gt;Data engineering is not shrinking. It's consolidating into a senior-heavy discipline that demands architectural thinking, governance awareness, and AI-infrastructure fluency. The tools change every 18 months. The problems don't change. Schema drift, late-arriving data, upstream teams breaking contracts without telling you. These are eternal.&lt;/p&gt;

&lt;p&gt;The ghost job epidemic will burn itself out eventually; it's expensive for companies too, even if they don't realize it yet. But the reclassification of entry-level to mid-level? That's structural. That's not going back.&lt;/p&gt;

&lt;p&gt;If you're grinding applications right now into what feels like a void: it's not you. Statistically, half of what you're applying to doesn't exist. That's not a personal failure. That's a broken system.&lt;/p&gt;

&lt;p&gt;What's the most obviously fake job posting you've come across, and how far into the process did you get before you realized it wasn't real?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>career</category>
      <category>beginners</category>
      <category>interview</category>
    </item>
    <item>
      <title>3 Staff Engineers Couldn't Pass This Single DE Job Posting</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Thu, 09 Jul 2026 10:08:18 +0000</pubDate>
      <link>https://dev.to/datadriven/3-staff-engineers-couldnt-pass-this-single-de-job-posting-4f0p</link>
      <guid>https://dev.to/datadriven/3-staff-engineers-couldnt-pass-this-single-de-job-posting-4f0p</guid>
      <description>&lt;p&gt;I sat down with two other staff-level data engineers last month. Between us: 40+ years in &lt;strong&gt;data engineering&lt;/strong&gt;, multiple FAANG stints, and enough &lt;strong&gt;interview&lt;/strong&gt; loops to fill a spreadsheet nobody asked for. We pulled up a single job posting from a Series C company with a real data team. Nothing exotic. Reasonable product, decent engineering culture, the kind of place you'd actually consider working.&lt;/p&gt;

&lt;p&gt;Not one of us met every requirement on the &lt;strong&gt;job description&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The posting wanted SQL, Python, Airflow, Snowflake, Kafka, Flink, dbt, Terraform, Kubernetes, and "LLM integration experience." Fifteen distinct technologies across four engineering disciplines. Salary: $140K to $170K. That's the same range these roles paid in 2022, when the ask was SQL, Python, Airflow, and a warehouse.&lt;/p&gt;

&lt;p&gt;Three staff engineers. 40+ combined years. Zero out of three qualified on paper.&lt;/p&gt;

&lt;p&gt;If that doesn't tell you the hiring process is broken, I don't know what does.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Impossible JD Is the New Normal
&lt;/h2&gt;

&lt;p&gt;In 2022, a typical data engineering posting listed four or five core tools. SQL. Python. An orchestrator (usually Airflow). A warehouse (usually Snowflake or BigQuery). Maybe Spark if the company was processing serious volume.&lt;/p&gt;

&lt;p&gt;In 2026, that same posting, at the same salary, now includes all of the above plus Kafka, Flink, dbt, Terraform, Kubernetes, and "LLM integration." LLM &lt;strong&gt;skills&lt;/strong&gt; in data engineering postings jumped from 3% to 12% in a single quarter. That's not gradual adoption; that's panic hiring.&lt;/p&gt;

&lt;p&gt;Here's the thing: these aren't complementary skills. They're separate &lt;strong&gt;career&lt;/strong&gt; tracks wearing a trench coat. A Databricks engineer working on distributed compute and Delta Lake optimization is not the same person as a Snowflake engineer designing warehouse concurrency patterns. SQL mastery and PySpark proficiency are taught in different phases of a career for a reason. Databricks engineers earn a $10K to $15K premium over Snowflake engineers at equivalent levels, not because Databricks is "better," but because production Spark plus distributed systems expertise is scarcer than SQL-first warehouse knowledge. These are fundamentally different learning curves, different architectures, different day-to-day work. Yet job descriptions list both as "required" like they're interchangeable checkboxes.&lt;/p&gt;

&lt;p&gt;Python appears in 70% of data engineer postings. SQL in 69%. But after that, the stack fragments: Spark at 38.7%, Snowflake at 29.2%, Kafka at 24%, Databricks at 16.8%. No two companies agree on what the stack actually is. So they list everything, hoping the perfect unicorn applies.&lt;/p&gt;

&lt;p&gt;The unicorn doesn't exist. And the engineers closest to it; the staff-level folks who understand exactly how complex this landscape is; they read that JD and close the tab.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;When experienced engineers see 14+ tools in one posting, many read it as "this company doesn't know what it needs," not "we're thorough."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Postings with 6 to 13 requirements receive about 30% more applicants than postings with 14+. That's not because the market lacks talent. It's because the people with the most experience and context are the ones most likely to self-select out. They know nobody does all of that. Junior candidates who don't know what they don't know are the ones clicking Apply.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Companies Keep Writing Fiction
&lt;/h2&gt;

&lt;p&gt;This isn't malice. It's organizational dysfunction.&lt;/p&gt;

&lt;p&gt;Most job descriptions aren't written by the person you'd actually report to. They're built by committee. The hiring manager adds SQL and Python. The VP adds Kubernetes and Terraform because "we're moving to cloud-native." The ML team bolts on LLM integration because they heard about it at a conference. The recruiter copy-pastes from a competitor's posting and sprinkles in whatever worked last time.&lt;/p&gt;

&lt;p&gt;The result: three different jobs inside one posting. Platform engineering. Streaming infrastructure. Analytics engineering. ML operations. Each one a distinct specialization with a different learning curve and a different interview loop. Nobody on the hiring committee realizes they've described four humans, not one. And nobody pushes back because adding requirements feels free.&lt;/p&gt;

&lt;p&gt;It's not free. Poorly written JDs are cited as the primary cause in over 50% of hiring failures. 44% of hiring managers can't fill their open roles, and the number is climbing. But here's the kicker: 94% of employers say skills-based hiring is more predictive of on-the-job success than resume screening, yet over half still screen against rigid checklists. They know the process is broken. They keep doing it anyway.&lt;/p&gt;

&lt;p&gt;The compensation tells the rest of the story. Data engineering salaries rose from about $113K to $153K over the past few years. That's a 35% bump. The role scope roughly doubled. You're being asked to learn twice as much, own twice as much, debug twice as much; for a 35% raise. The economics don't work, and experienced engineers can see that from the posting alone.&lt;/p&gt;

&lt;p&gt;And let's talk about the 40% of tech companies that posted ghost jobs in the past year. Nearly half of visible data engineering roles on LinkedIn aren't actual open positions. You're grinding through a 15-tool requirement list for a role that may not even exist. 62% of hiring managers admit their AI screening tools reject qualified candidates who don't match algorithmic patterns. So even if you apply, the ATS might bury you before a human reads your name. The filter is broken at every level.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Job Actually Requires on Day One
&lt;/h2&gt;

&lt;p&gt;Here's what happens your first week at any of these companies. Nobody asks you to configure a Flink cluster, deploy a model to Kubernetes, and integrate an LLM into the data pipeline before lunch. What actually happens: someone points you at a broken pipeline, and you figure out why it's broken.&lt;/p&gt;

&lt;p&gt;SQL and Python still carry the game. SQL appears in 94% of interview loops for data engineering roles, and SQL plus Python together get candidates through roughly 60% of real interviews. The other eight tools on the posting? Most teams deploy three, maybe four, in production. The rest is aspirational.&lt;/p&gt;

&lt;p&gt;The actual job is less "architect a real-time streaming platform" and more "figure out why this pipeline silently dropped 2M rows last Tuesday and make sure it never happens again." Production work (debugging, incident response, observability) comprises around 60% of what data engineers do day to day. It appears in approximately 0% of job descriptions. Nobody interviews for that skill. They interview for Spark API trivia and LeetCode mediums. These are measuring different things entirely.&lt;/p&gt;

&lt;p&gt;77% of data engineers report heavier workloads in 2026 despite AI tools that were supposed to lighten the load. Engineers now spend 37% of their time on AI-related projects, up from 19% two years ago. AI didn't replace work; it created new categories of work on top of the existing ones. The tooling promise of "do more with less" turned into "do more with more, and also learn this new thing by Friday."&lt;/p&gt;

&lt;p&gt;So what actually matters? &lt;strong&gt;Data modeling&lt;/strong&gt;. Query optimization. Understanding why things break, not just how to set things up when they're working. Pipeline architecture, not system design (data engineers don't care about load balancers and reverse proxies). The concepts that transfer across every tool, every warehouse, every orchestrator. Concepts transfer; tool knowledge doesn't. That has always been the thesis, and bloated job descriptions haven't changed it one bit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 70% Rule and How to Play the Game
&lt;/h2&gt;

&lt;p&gt;Here's the practical advice. Apply at 70%. Not 70% of every random tool listed; 70% of the must-haves. SQL, Python, one orchestrator, one warehouse. If you've got those plus a track record of shipping and debugging production pipelines, you're in the conversation.&lt;/p&gt;

&lt;p&gt;42% of applicants don't meet posted requirements. Companies fill those roles anyway. Hiring managers often know they won't get someone who ticks every box; the posting is a negotiation anchor, not a contract. 81% of U.S. employers have adopted skills-based hiring, up from 57% in 2022. The JD says 15 tools; the interview tests three.&lt;/p&gt;

&lt;p&gt;Stop trying to learn Kafka and Flink and Terraform and Kubernetes simultaneously. That's tool chasing, and it's a trap. Double down on SQL, Python, and data modeling. Get your reps in on the stuff that actually gets asked. If you're sharpening the Python side, we put together python interview questions on datadriven specifically around patterns that show up in real loops, not obscure trivia nobody's ever tested on in an actual interview.&lt;/p&gt;

&lt;p&gt;Optimize your resume for the three or four tools the team actually uses, not the 15 they listed. Don't match fiction with fiction. If a job description has more than six or seven distinct tools in the requirements, odds are it was committee-built and nobody on the actual team uses all of them. That's not a signal to disqualify yourself; it's a signal that the company will hire for the core and train for the periphery.&lt;/p&gt;

&lt;p&gt;Strip back the scope anxiety. Senior and staff titles are converging on the same JD with $10K salary deltas. On paper, there's nowhere to grow. But in practice, the engineer who can debug a production incident, model a clean schema, and explain their design decisions under pressure is the one who gets hired and promoted. That hasn't changed in 15 years of data engineering, and it won't change because someone added "LLM integration" to a posting.&lt;/p&gt;

&lt;p&gt;Three staff engineers, 40+ combined years, and none of us passed a single job description. That should liberate you, not discourage you. The posting is fiction. Your skills are real. Know the difference, apply anyway, and let the interview sort it out.&lt;/p&gt;

&lt;p&gt;What's the most absurd set of requirements you've seen crammed into a single data engineering posting?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>career</category>
      <category>interview</category>
      <category>beginners</category>
    </item>
    <item>
      <title>AI Is Cutting DE Jobs. It's Also Creating Them. Here's the Map.</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 07 Jul 2026 10:07:35 +0000</pubDate>
      <link>https://dev.to/datadriven/ai-is-cutting-de-jobs-its-also-creating-them-heres-the-map-2pn1</link>
      <guid>https://dev.to/datadriven/ai-is-cutting-de-jobs-its-also-creating-them-heres-the-map-2pn1</guid>
      <description>&lt;p&gt;I watched a friend get laid off from Meta's analytics org on a Tuesday. By Thursday, Meta had posted six new roles in their "Applied AI Engineering" unit that required almost identical pipeline experience. Different title. Different budget line. Same category of work, but pointed at LLMs instead of dashboards.&lt;/p&gt;

&lt;p&gt;That's the game right now. And most people don't have the map.&lt;/p&gt;

&lt;h2&gt;
  
  
  52,050 Cuts. 23% Hiring Growth. Both Are True.
&lt;/h2&gt;

&lt;p&gt;Q1 2026 produced &lt;strong&gt;52,050 tech job cuts&lt;/strong&gt;, the highest Q1 total since 2023; a 40% jump over Q1 2025. Of those layoff events, 56% explicitly cited &lt;strong&gt;AI&lt;/strong&gt; as the reason. Meta cut 8,000 people. Intuit eliminated 3,000 (17% of their workforce). PayPal axed 4,760 (20%). Salesforce trimmed under 1,000 from data analytics and product management.&lt;/p&gt;

&lt;p&gt;And yet: &lt;strong&gt;data engineering&lt;/strong&gt; hiring is up 23% year-over-year, with roughly 260,000 open US positions projected for 2026. The data engineering market hit $105 billion this year and is projected to reach $213 billion by 2031.&lt;/p&gt;

&lt;p&gt;These aren't contradictory numbers. They're describing two different jobs that happen to share a title.&lt;/p&gt;

&lt;p&gt;The version of data engineering that was "connect source A to warehouse B using tool C" is getting automated. The version that involves distributed systems, governance, cost optimization, and AI infrastructure is in acute shortage. If you're a mid-level engineer watching both headlines and feeling whiplash, it's because you're standing on the fault line between a role that's dying and one that's exploding. The World Economic Forum forecasts 100% growth in big data specialist demand through 2030. Meanwhile, 63% of employers say they can't find qualified candidates. If supply were truly saturated from all these &lt;strong&gt;layoffs&lt;/strong&gt;, hiring timelines would compress. They haven't. Time-to-hire for data engineers is stuck at 60 to 90 days in enterprise settings.&lt;/p&gt;

&lt;p&gt;The paradox resolves when you stop thinking of "data engineer" as one job.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cut-and-Redirect Pattern
&lt;/h2&gt;

&lt;p&gt;Here's what Meta actually did: cut 8,000 employees, then simultaneously moved 7,000 workers into AI-focused units called "Applied AI Engineering" and "Agent Transformation Accelerator." That's not a net reduction; it's a budget reallocation with human casualties.&lt;/p&gt;

&lt;p&gt;Atlassian did the same thing at smaller scale: 1,600 jobs eliminated, 800 AI roles posted within weeks. PayPal cut 20% of its workforce while immediately posting for AI infrastructure and ML pipeline engineers. The pattern is so consistent it deserves its own name.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Companies aren't reducing data headcount. They're killing one version of the role and resurrecting another, and the six-month gap between the cut and the rehire is where careers go to die.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here's the number that should make you uncomfortable: &lt;strong&gt;52.1% of companies making AI-driven layoffs rehired for nearly identical roles within six months&lt;/strong&gt;. Not identical titles, but identical skill categories. The press release says "AI eliminated these positions." The LinkedIn posting six months later says "seeking experienced data engineers for AI platform." Same budget. Same org chart slot. Different words.&lt;/p&gt;

&lt;p&gt;And 55% of business leaders who pulled the trigger on AI-driven cuts now regret the decision, after discovering that AI handles 94% of routine tasks but falls apart on judgment calls. IBM and Ford have been quietly rehiring. Nobody writes a press release about that part.&lt;/p&gt;

&lt;p&gt;Anthropic, OpenAI, Cohere, and Mistral hired 6,200 people combined in H1 2026. Record pace. The talent has somewhere to go; the problem is that most displaced engineers don't have the specific skills those roles demand. Roughly 40% of displaced workers land in mid-market companies, 25% become contractors, and only about 15% actually change specialties. The pipeline from "laid off ETL engineer" to "hired AI infrastructure engineer" has a massive leak in the middle.&lt;/p&gt;

&lt;h2&gt;
  
  
  What AI Actually Made Cheap
&lt;/h2&gt;

&lt;p&gt;Let's be specific about what got commoditized, because vague panic helps nobody.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Basic SQL generation.&lt;/strong&gt; Text-to-SQL tools now hit about 78% execution accuracy on zero-shot queries. That sounds impressive until you realize one in five queries is silently wrong: hallucinated columns, faulty joins, dropped WHERE clauses, missing tenant scoping. The query runs. It returns results. The results are wrong. Nobody notices for weeks. I've seen this movie before; it used to star junior analysts instead of LLMs. The failure mode is identical; the scale is larger.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Boilerplate ETL scripting.&lt;/strong&gt; Source-to-target mapping, schema detection, field mapping. Databricks reported that 80% of new databases on their platform are now created by AI agents, up from 30% a year ago. The plumbing got automated. If your entire job was plumbing, you should be worried.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AutoML patterns.&lt;/strong&gt; Hyperparameter search, basic feature engineering, architecture exploration. Cycles that took weeks now take hours. But AutoML doesn't recognize when the problem framing itself is wrong. It optimizes within the frame you give it; it can't tell you the frame is garbage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dashboard maintenance and ad-hoc reporting.&lt;/strong&gt; Data analyst job openings fell roughly 40% from peak. The "pull this data for me" function is collapsing into self-service AI tooling. Microsoft Copilot, Tableau AI, and their cousins are eating this alive.&lt;/p&gt;

&lt;p&gt;Here's what didn't get commoditized: schema evolution decisions. Data contracts. Lineage governance. Cost optimization at scale. Figuring out why your pipeline silently dropped 2M rows last Tuesday and making sure it never happens again. The judgment layer is intact. The mechanical layer is not.&lt;/p&gt;

&lt;p&gt;This is why the old advice still holds: concepts transfer across tools; tool knowledge doesn't transfer across concepts. The engineers who spent their &lt;strong&gt;career&lt;/strong&gt; learning "how Airflow works" are in trouble. The ones who spent it learning "how to design systems that don't fail silently" are getting promoted.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Map: What's Actually Getting Hired
&lt;/h2&gt;

&lt;p&gt;If you're navigating this transition, you need specifics, not vibes. Here's where the &lt;strong&gt;hiring&lt;/strong&gt; is happening, and what it pays.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MLOps and AI infrastructure.&lt;/strong&gt; Salaries range $130K to $257K, with LLM deployment experience consistently pushing past $200K. AI infrastructure engineer salaries jumped 15 to 30% in H1 2026, averaging $320,000 in San Francisco. This is the hottest category by far, and it's not close.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Streaming and real-time systems.&lt;/strong&gt; $114K to $137K average. I still think streaming is overrated for most companies (most of y'all don't need it), but the ones that need it are paying well and hiring fast. Confluent cut 800 employees; Databricks had 840 open requisitions the same month and actively recruited from that pool. That's not market equilibrium; that's targeted talent consolidation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data governance.&lt;/strong&gt; Over 6,500 open governance roles in the US, averaging $113K annually. GDPR pressure, AI governance requirements, and the realization that you can't ship AI products on dirty data are driving this. Not the sexiest work. Pays reliably.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FinOps and cloud cost optimization.&lt;/strong&gt; $98K to $167K for cloud FinOps roles. Every company that went all-in on cloud compute is now discovering the bill. They need people who understand both the data and the economics. If you consider the cost of running unoptimized pipelines versus the cost of the engineer's time optimizing them, the economics clearly favor hiring the engineer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's conspicuously absent from the hiring surge:&lt;/strong&gt; junior-to-mid-level generalist data engineers. The 23% growth is concentrated at the senior-specialist tier. Engineers under 30 saw the greatest decline. Job descriptions are collapsing three roles into one: platform engineering, ML pipeline support, and governance orchestration. That's scope creep masquerading as demand.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Comp Bifurcation
&lt;/h2&gt;

&lt;p&gt;Now here's the part nobody talks about honestly. The AI salary premium over traditional DE roles widened from 25% in 2024 to 56% in 2026. A traditional ETL developer averages $145K. An AI/ML engineer averages $173K to $193K base. LLM specialists command $220K to $280K. Frontier lab researchers clear $600K.&lt;/p&gt;

&lt;p&gt;FAANG senior data engineer total comp sits at $250K to $350K, with stock composing more than half the package. Non-FAANG senior comp: $105K to $175K. Stripe pays about $210K for senior DE, which sounds great until you realize it's still 33% below Google's floor. Even at the 95th percentile of non-FAANG, you're hitting the 25th percentile at Google.&lt;/p&gt;

&lt;p&gt;The real dividing line isn't FAANG versus startups. It's equity-heavy jobs versus equityless jobs. A data engineer at Meta L5 makes roughly $350K all-in with about $250K from RSUs vesting linearly. A data engineer at a healthcare company makes $140K all-cash with no equity. One of these paths builds generational wealth. The other one consumes it.&lt;/p&gt;

&lt;p&gt;A $150K ETL architect doesn't become a $320K AI infrastructure engineer without 12 to 18 months of intentional retraining. The skills are orthogonal, not adjacent. That's the structural trap nobody in HR wants to acknowledge.&lt;/p&gt;

&lt;p&gt;So what do you actually do about it? Stop treating this like a tool problem. Data modeling, distributed systems thinking, cost optimization, governance: these are the concepts that transfer. The specific AI tools will change in 18 months. They always do. I've been through three waves of "data engineering is getting automated away." Still here. Still employed. Still debugging the same categories of problems, just with fancier abstractions on top.&lt;/p&gt;

&lt;p&gt;Interviewing is a separate skill from the actual job, and the bar has shifted. If you want to pressure-test where you stand, that's exactly why we built out our &lt;a href="https://datadriven.io" rel="noopener noreferrer"&gt;datadriven.io data engineer interview questions&lt;/a&gt; around architectural thinking and system design rather than tool trivia; the roles getting hired today require a fundamentally different interview prep than the ones getting cut.&lt;/p&gt;

&lt;p&gt;The engineers who survive this transition won't be the ones who learned the most AI tools. They'll be the ones who understood the problems well enough that the tools didn't matter. Junior engineers worry about which tool to learn. Senior engineers worry about which problems to solve. Staff engineers worry about which problems to prevent.&lt;/p&gt;

&lt;p&gt;Figure out which category you're in. Then move up.&lt;/p&gt;

&lt;p&gt;What's the biggest shift you've noticed in DE job descriptions over the last year? I'm curious whether the bifurcation looks the same from inside healthcare, fintech, and pure tech, or if it's playing out differently by industry.&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>career</category>
      <category>python</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Data Engineer LeetCode: What to Grind and What to Skip</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Fri, 03 Jul 2026 17:06:34 +0000</pubDate>
      <link>https://dev.to/datadriven/data-engineer-leetcode-what-to-grind-and-what-to-skip-4emk</link>
      <guid>https://dev.to/datadriven/data-engineer-leetcode-what-to-grind-and-what-to-skip-4emk</guid>
      <description>&lt;p&gt;I spent about 80 hours grinding LeetCode before my first FAANG &lt;strong&gt;data engineering&lt;/strong&gt; loop. Binary trees, dynamic programming, graph traversal. I could reverse a linked list in my sleep. Then I walked into the interview and got asked to deduplicate a fact table with late-arriving records, design a pipeline for slowly changing dimensions, and write a window function I could have done in 10 minutes if I hadn't been so sleep-deprived from memorizing Dijkstra's algorithm the night before.&lt;/p&gt;

&lt;p&gt;I bombed it. Not because I wasn't prepared. Because I prepared for the wrong test.&lt;/p&gt;

&lt;p&gt;That was years ago, and the gap between what &lt;strong&gt;LeetCode&lt;/strong&gt; tests and what data engineering &lt;strong&gt;interviews&lt;/strong&gt; actually screen for has only gotten wider. In 2026, candidates are still burning hundreds of hours on problem types that virtually never surface in DE loops, while the skills that actually separate hire from no-hire get treated as afterthoughts. SQL fluency, data-manipulation &lt;strong&gt;Python&lt;/strong&gt;, pipeline design thinking. That's where offers come from. Not from memorizing Dijkstra's.&lt;/p&gt;

&lt;p&gt;Let me save you some time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The LeetCode Mismatch Nobody Talks About Honestly
&lt;/h2&gt;

&lt;p&gt;Here's the math that should make you angry: there are over 3,000 problems on LeetCode. The vast majority test binary tree traversal, dynamic programming, graph algorithms, and backtracking. Data engineering interviews rarely touch any of those categories.&lt;/p&gt;

&lt;p&gt;62% of organizations prohibit AI use in technical interviews, while 76% of actual data engineering work is now enhanced by AI tools. So the interview tests skills the job doesn't use, in conditions the job doesn't impose. If an AI can spit out a clean solution to a medium LC problem, what does asking that problem actually tell anyone about the candidate? The signal has always been thin. Now it's basically noise.&lt;/p&gt;

&lt;p&gt;The industry knows this. Between 2023 and 2026, the DE role shifted from "batch ETL plumber" to a blend of real-time architecture, cloud cost optimization, metadata governance, and platform engineering. None of that correlates to your ability to implement a trie. Companies are starting to act on it: Airbnb's loop dropped dedicated coding puzzle stages in favor of pipeline design rounds. Meta replaced traditional LeetCode screens with staged CodeSignal scenarios. Google now hands candidates multi-file codebases for refactoring instead of isolated algorithm puzzles.&lt;/p&gt;

&lt;p&gt;But candidates? Still grinding binary trees at 2am.&lt;/p&gt;

&lt;p&gt;The research is clear: 35 to 50 problems is sufficient for most data engineering roles. 10 to 15 easy, 20 to 25 medium, 5 to 10 hard. That's it. Skip trees, linked lists, graphs, and backtracking entirely unless a specific company tells you otherwise. Stick to arrays, hash maps, string manipulation, and sliding windows. These are the patterns that actually transfer to data work; the rest is noise you're studying to feel productive.&lt;/p&gt;

&lt;p&gt;Dynamic programming is nearly useless for DE work but still appears in prep checklists. Most DP problems aren't applicable in real-world settings, yet candidates grind them out of habit, wasting 20 to 40 hours on dead-end prep. I know because I did exactly that. I memorized the knapsack problem. Never once used it. Not once.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The best data engineers often aren't strong at algorithm puzzles. The reverse is also true. Stop optimizing for the wrong metric.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  SQL Is the Real Interview Gate
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;SQL&lt;/strong&gt; appears in 69 to 79% of data engineer job postings. It shows up in 85% of full interview loops. Window functions appear on roughly 80% of technical screens. If you can't write &lt;code&gt;ROW_NUMBER()&lt;/code&gt; or a rolling average with &lt;code&gt;OVER(PARTITION BY ...)&lt;/code&gt; cold, you're going to struggle at any data-adjacent role. These aren't advanced anymore. They're table stakes.&lt;/p&gt;

&lt;p&gt;The patterns that actually gate candidates are narrower than most people assume.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Window functions.&lt;/strong&gt; &lt;code&gt;ROW_NUMBER()&lt;/code&gt; vs. &lt;code&gt;RANK()&lt;/code&gt; vs. &lt;code&gt;DENSE_RANK()&lt;/code&gt;. Getting the wrong one when ties exist cascades into broken analytics. &lt;code&gt;LAG&lt;/code&gt; and &lt;code&gt;LEAD&lt;/code&gt; for sessionization and gap detection. This is the single skill that separates junior from intermediate in the eyes of most interviewers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CTEs.&lt;/strong&gt; Interviewers no longer accept nested subqueries as a clean approach. Break your logic into named steps. If your query reads like a paragraph instead of a matryoshka doll, you're already ahead of 60% of candidates. When I tried to submit my first bit of SQL to my code repository, the response I received must have been longer than the code submitted. I didn't understand the value of readability. I do now.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Join-grain awareness.&lt;/strong&gt; One in three real SQL rounds opens with a problem where candidates inflate revenue by 3x because they joined at the wrong grain. The number one failure mode isn't syntax; it's not understanding the cardinality of the relationship before writing the JOIN.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deduplication.&lt;/strong&gt; If you're slapping &lt;code&gt;DISTINCT&lt;/code&gt; on a query to hide a problem instead of solving it, the interviewer noticed. Use &lt;code&gt;ROW_NUMBER()&lt;/code&gt; to deduplicate on a composite key. Know when your data has duplicates because of the source vs. because of your join.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NULL behavior.&lt;/strong&gt; This one is a silent killer. A single NULL in a subquery makes &lt;code&gt;NOT IN&lt;/code&gt; return zero rows. Not the filtered set you expected; zero rows. This defeats roughly one in five candidates. Use &lt;code&gt;NOT EXISTS&lt;/code&gt; instead. It handles NULLs correctly and it's what your interviewer wants to see.&lt;/p&gt;

&lt;p&gt;Here's what most candidates miss: clarification beats speed. If the question says "find the latest order," does "latest" mean by timestamp or by ID? Candidates who jump straight to coding burn 20 to 30 minutes solving the wrong problem. The ones who ask two questions first finish in 10.&lt;/p&gt;

&lt;p&gt;Phone screens use 2 to 3 conceptual questions. On-sites use 4 to 6 hands-on problems. No binary trees. No DP. Just window functions, CTEs, grain, and dedup. That's the SQL interview study plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  Python Rounds Want a Data Brain, Not an Algo Brain
&lt;/h2&gt;

&lt;p&gt;Python appears in 74% of data engineer job postings and is required for senior roles. But the Python that matters in a DE interview is completely different from what SWE screens test.&lt;/p&gt;

&lt;p&gt;Uber's data engineer screen asks candidates to transform transaction datasets and calculate custom metrics using Pandas. Stripe emphasizes clean, efficient Python with a focus on data structures and SQL first, then scalable pipeline design. No graph traversal. No dynamic programming. The actual screening questions look like work you'd do on the job: parse a messy log file, deduplicate records on a composite key, sessionize an event stream, walk a nested JSON structure.&lt;/p&gt;

&lt;p&gt;Most candidates don't fail because they can't write Python. They fail on the one malformed row in a file of ten million. The interview tests whether you think about validation, error handling for bad data, and what happens when a field is missing or the wrong type. Can you quarantine bad records and keep the pipeline running? Or does your code silently drop 40% of rows because you assumed clean input?&lt;/p&gt;

&lt;p&gt;I've seen candidates crush 100 LeetCode problems and then stumble on a composite-key deduplication task. They memorized algorithms; they never learned to think about data. The skills that gate most DE candidates (window functions, idiomatic Pandas, JSON schema handling, idempotent upsert logic) are entirely absent from the "hard" LeetCode catalog.&lt;/p&gt;

&lt;p&gt;Here's the counterintuitive thing: SQL fluency now outweighs Python sophistication for most loops. Clean, efficient SQL signals maturity within 10 minutes. Advanced Python (decorators, asyncio, metaprogramming) rarely surfaces. If you have limited prep time, put 40% into SQL, 30% into Python data manipulation, 20% into system design, and 10% into behavioral. That split comes from analyzing actual interview loops, not from some prep course syllabus.&lt;/p&gt;

&lt;p&gt;For the PySpark and data-manipulation side specifically, we put together targeted drills for exactly this kind of prep; datadriven.io is good for pyspark practice and the pattern-matching that actually transfers to live screens, so you're not wasting reps on algorithm trivia.&lt;/p&gt;

&lt;h2&gt;
  
  
  System Design Replaced the Algorithm Round
&lt;/h2&gt;

&lt;p&gt;Here's the shift that caught everyone off guard: system design expectations moved down the seniority ladder. What used to screen only senior engineers now appears at mid-level interviews. Airbnb's final loop includes 5 to 7 rounds with 1 to 2 system design rounds heavily weighted for leveling decisions. At Meta, it's the highest-weighted round in modern senior loops. At Databricks, you're designing real-time fraud detection using Spark Structured Streaming, Kafka, and Delta Lake.&lt;/p&gt;

&lt;p&gt;But DE system design isn't SWE system design. Strip back the "system design for software engineers" mentality. You don't need to hand-roll a message broker or explain Paxos consensus. You need to reason about slowly changing dimensions, schema drift handling, idempotent writes, and which warehouse suits the cardinality and latency profile of the problem. This knowledge comes from building pipelines, not reading papers.&lt;/p&gt;

&lt;p&gt;71% of engineering leaders report AI is making it harder to assess candidates, which is accelerating the format shift. The old playbook (memorize algorithms, speed-run solutions, pray the interviewer asks something you've seen) is dying. The new playbook rewards systems thinking, cost reasoning, and the ability to navigate ambiguity. Communication, narration, and reasoning under pressure are now the primary differentiators; not whether you can implement quicksort from memory.&lt;/p&gt;

&lt;p&gt;The candidates who get offers aren't the ones with perfect algorithm solutions. They're the ones who ask the right questions before writing anything. "What's the expected data volume? How fresh does the downstream consumer need it? What happens when the upstream schema changes without warning?" A candidate who stumbles on a medium-easy coding problem but reasons clearly through pipeline architecture gets hired over the person who solves the algorithm perfectly but can't articulate a single tradeoff.&lt;/p&gt;

&lt;p&gt;Data modeling is the quiet kingmaker here. It's the most important part of any data engineering interview, and if you nail the technical coding but stumble on modeling, you likely won't get the offer. Getting the model wrong upstream means everything downstream is pain. I've watched people with 10 YOE get downleveled because they couldn't articulate schema design decisions under pressure. The interview is a different skill than the job.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Junior engineers worry about which tool to learn. Senior engineers worry about which problems to solve. Staff engineers worry about which problems to prevent.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The data engineering interview in 2026 tests whether you can think, not whether you can memorize. DP, binary trees, and graph algorithms burned hours of my life that I'll never get back. Window functions, CTE fluency, data-manipulation Python, and pipeline design thinking are what got me hired. Repeatedly.&lt;/p&gt;

&lt;p&gt;If you're in the middle of a job search right now, reclaim your prep time. Drop the hard LeetCode grind, double down on SQL patterns and pipeline architecture, and treat interviewing like the separate skill it is. The tools change every 18 months. The problems don't. Schema drift, late-arriving data, upstream teams breaking contracts without telling you. These are eternal. Study the eternal stuff.&lt;/p&gt;

&lt;p&gt;What's the single interview question that caught you most off-guard in a DE loop? The ones nobody warned you about are the ones worth sharing.&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>interview</category>
      <category>sql</category>
      <category>career</category>
    </item>
    <item>
      <title>DE Interviews Dropped DSA. The Replacement Is a Mess.</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 30 Jun 2026 10:07:17 +0000</pubDate>
      <link>https://dev.to/datadriven/de-interviews-dropped-dsa-the-replacement-is-a-mess-5g8o</link>
      <guid>https://dev.to/datadriven/de-interviews-dropped-dsa-the-replacement-is-a-mess-5g8o</guid>
      <description>&lt;p&gt;I did somewhere around 20 interview loops in a single job search. Some tested me on binary trees. Some tested me on pipeline architecture. One company asked me to build a full data warehouse from scratch in a take-home, then ghosted me. Another had me whiteboard a Spark optimization problem, then admitted in the debrief that they don't use Spark. The only consistent thing about &lt;strong&gt;data engineering&lt;/strong&gt; interviews in 2025 and 2026 is that nothing is consistent.&lt;/p&gt;

&lt;p&gt;The industry quietly dropped &lt;strong&gt;DSA&lt;/strong&gt; rounds from DE hiring pipelines. You'd think that would be good news. Data engineers don't implement Dijkstra's algorithm at work. We debug pipelines that silently drop records, negotiate schema contracts with upstream teams who don't know we exist, and model data so finance can build board decks without calling us at midnight. &lt;strong&gt;LeetCode&lt;/strong&gt; mediums were always an imported ritual from software engineering, never validated for our work.&lt;/p&gt;

&lt;p&gt;But here's the thing nobody warned you about: the replacement is worse.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why DE Interviews Ever Tested DSA
&lt;/h2&gt;

&lt;p&gt;Data engineering solidified as a distinct role somewhere around 2015 to 2020. That's recent. Software engineering has 40+ years of structured &lt;strong&gt;interview&lt;/strong&gt; evolution. When companies needed to hire DEs, they didn't design DE-specific screening. They copy-pasted the SWE playbook. LeetCode reached a million users within two years of launch, and suddenly every technical role in tech was getting filtered through binary tree traversals and dynamic programming problems.&lt;/p&gt;

&lt;p&gt;The logic was lazy but understandable: "Smart people can solve algorithm problems. We want smart people. Therefore, algorithm problems." Nobody asked whether the signal was relevant. 78% of developers report that interview assessments don't match real-world job work, and 56% say algorithm questions are "not useful for their jobs." For data engineers, that number should be higher. We rarely write complex algorithms from scratch; we use pre-built libraries and frameworks. The skill is knowing which tool to reach for and how to model the data correctly, not implementing a red-black tree.&lt;/p&gt;

&lt;p&gt;Inertia kept DSA in place for years. Not evidence. Algorithm interviews are easy to design, easy to administer, and easy to score. They have clear right answers. They let hiring committees feel rigorous without doing the hard work of defining what "good DE" actually looks like. Seven out of ten companies were still screening data engineers the same way they did in 2022, even as the role itself morphed from "batch ETL plumber" to something combining real-time architecture, cloud cost optimization, metadata governance, and AI integration.&lt;/p&gt;

&lt;p&gt;Then AI made the whole thing absurd. If Claude can solve a medium LC problem in seconds, what does asking it tell you about the candidate? Companies like Meta, Google, Canva, and Shopify started permitting AI use in live technical sessions. Canva replaced their "Computer Science Fundamentals" round with "AI-Assisted Coding" in mid-2025. The premise that raw algorithmic ability was being measured collapsed overnight.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;DSA was never the right test for data engineering. It was the convenient one. And when convenience stopped working, nobody had a Plan B.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What Companies Replaced the LeetCode Round With (And Why It's a Mess)
&lt;/h2&gt;

&lt;p&gt;Here's where the story gets ugly. Companies dropped DSA and replaced it with whatever their hiring manager felt like that quarter. There is no standard. There is no consensus. There is barely even a pattern.&lt;/p&gt;

&lt;p&gt;Some companies pivoted to &lt;strong&gt;system design&lt;/strong&gt; rounds. Fine in theory; pipeline architecture is genuinely relevant to the job. But system design has no rubric. No "correct" answer. Every interviewer has a different opinion on whether you should optimize for cost, latency, or data freshness. At least with LeetCode, everyone agreed what a good answer looked like. System design is subjective hell, and candidates walk out of those rounds with zero idea whether they passed.&lt;/p&gt;

&lt;p&gt;Other companies went all-in on take-home projects. These started as 2 to 3 hour exercises and ballooned into 10 to 20 hour ordeals: full pipeline implementations, multi-source data modeling, documentation, testing, and presentation follow-ups. That's not an interview; that's free proof-of-concept work. And the rejection email is still a template with no feedback.&lt;/p&gt;

&lt;p&gt;Here's the part that should make you angry: take-homes are worse for equity than live coding. A structured live interview can be leveled with a rubric. A 20-hour take-home is a time tax that penalizes caregivers, people working second jobs, and anyone without unlimited buffer hours. Companies adopted take-homes thinking they were more fair. The irony is thick.&lt;/p&gt;

&lt;p&gt;The numbers paint the chaos clearly. SQL shows up in 85% of loops. System design in 65%. Python in 70%. &lt;strong&gt;Data modeling&lt;/strong&gt; in 55%. Take-homes in about 25%. Enterprise hiring timelines now stretch 60 to 90 days with 5 to 7 rounds, while the best candidates leave the market in 10 to 14 days. The process is optimized for companies feeling thorough, not for actually hiring.&lt;/p&gt;

&lt;p&gt;And the single most important DE skill, data modeling, is missing from two-thirds of interview loops. Only about 33% include a dedicated data modeling round. You can nail the coding, ace the system design whiteboard, and still lose the offer because you stumbled on modeling. But nobody told you that was the real test, because it's buried inside other rounds instead of being evaluated explicitly.&lt;/p&gt;

&lt;p&gt;The role definition chaos explains the interview chaos. Between 2023 and 2026, the industry moved DE from "batch ETL" to a role combining real-time architecture, AI pipelines, metadata governance, and cost optimization. Companies testing SQL plus system design plus AI-assisted builds are simultaneously hiring for three different job titles. No wonder candidates prep for interviews that don't exist once they're hired.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI Broke the Fallback Too
&lt;/h2&gt;

&lt;p&gt;Take-homes were supposed to be the safe harbor. Let candidates work in their natural environment, at their own pace, showing real engineering judgment. That premise died when LLMs got good enough to build a passable data pipeline in 20 minutes.&lt;/p&gt;

&lt;p&gt;The numbers are damning. 80% of candidates use LLMs on take-home tests despite explicit prohibition. AI-assisted cheating adoption jumped from 15% in June 2025 to 35% by December 2025, and by late 2026 it's headed past 50%. In unproctored take-home formats, estimated fraud rates sit between 60% and 80%. And 61% of cheaters scored above passing thresholds without detection.&lt;/p&gt;

&lt;p&gt;Companies responded with policies that are pure theater. 64% of companies attempt to ban AI in interviews, but there's zero correlation between having an explicit no-AI policy and lower cheating rates. The 80% use-despite-ban number tells you everything: candidates view bans as unenforceable, because they are. Tools like Cluely and Final Round AI cost $20 to $50 a month and feed answers via invisible screen overlays. Keystroke dynamics, perplexity scoring, gaze tracking; every detection method has documented, production bypasses.&lt;/p&gt;

&lt;p&gt;Greenhouse's June 2026 report found 80% of US candidates say employer AI policies are vague, rare, or completely absent. Companies blame candidates for guessing; candidates blame companies for silence. Meanwhile, 41% of companies now require a hybrid model: asynchronous take-home plus live defense session, specifically because unproctored take-homes alone produce no reliable signal. That's not marketed as "AI-proofing." It's just the new floor.&lt;/p&gt;

&lt;p&gt;The job market itself is training candidates to cheat. When 30% of candidates drop out of hiring processes after discovering AI-led screening, and take-homes are the fallback, the lesson is clear: assume no human will review this carefully and act accordingly. The policy vacuum creates the behavior.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;64% ban LLMs. 80% use them anyway. That's not a policy; it's a suggestion.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What to Actually Prep for When the Interview Has No Standard
&lt;/h2&gt;

&lt;p&gt;I'll be blunt: the chaos is the prep. You can't study for a standardized loop because there isn't one. But you can build a stack of skills that covers the majority of what companies actually test, regardless of which format they chose this quarter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SQL is still the universal filter.&lt;/strong&gt; It appears in 85% of loops. Not "write a SELECT statement" SQL; deep SQL. Window functions, CTEs, query optimization, understanding execution plans. This is the one skill that transfers across every company and every format, which is exactly why we made sure DataDriven is good for sql interview practice that reflects what companies actually ask rather than textbook exercises disconnected from real pipelines.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data modeling is the hidden pass/fail.&lt;/strong&gt; It's only an explicit round in a third of loops, but it's the implicit evaluation in every system design and take-home. Getting the model wrong upstream means everything downstream is pain. Practice designing schemas for real business scenarios: slowly changing dimensions, event streams, aggregation trade-offs. If you can't explain why you chose a grain, you're not ready.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;System design means pipeline architecture, not load balancers.&lt;/strong&gt; Strip back the "system design for software engineers" mentality. DEs don't care about reverse proxies. You need to design ingestion, transformation, serving layers, and failure handling. Practice narrating your design decisions out loud. The signal in system design rounds is communication under ambiguity, not arriving at the "right" architecture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Learn to debug, not just build.&lt;/strong&gt; The actual job is less "write a DAG" and more "figure out why this pipeline silently dropped 2M rows last Tuesday." Schema drift, late-arriving data, upstream teams breaking contracts without telling you; these are eternal. No interview tests for this explicitly, but candidates who can reason about failure modes stand out in every format.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat interviewing as a separate skill from the job.&lt;/strong&gt; I've watched people with 10 YOE get downleveled because they couldn't articulate system design decisions under pressure. The &lt;strong&gt;career&lt;/strong&gt; progression from senior to staff depends on your ability to communicate tradeoffs, not just execute them. Practice talking through your work. Record yourself explaining a pipeline you built. It feels stupid. It works.&lt;/p&gt;

&lt;p&gt;The LeetCode era had one advantage: predictability. You knew the game, you ground the reps, you played to win. The post-DSA era took that away without replacing it with anything coherent. That's frustrating. It's also the reality.&lt;/p&gt;

&lt;p&gt;The companies that figure this out first will attract the best talent. The ones still running 7-round loops with no rubric and 20-hour take-homes will keep wondering why their offers get declined.&lt;/p&gt;

&lt;p&gt;What's the most absurd interview format you've encountered since companies started dropping DSA rounds?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>interview</category>
      <category>career</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Data Engineer Salaries in 2026: The Numbers Are Lying</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Thu, 25 Jun 2026 10:07:39 +0000</pubDate>
      <link>https://dev.to/datadriven/data-engineer-salaries-in-2026-the-numbers-are-lying-5877</link>
      <guid>https://dev.to/datadriven/data-engineer-salaries-in-2026-the-numbers-are-lying-5877</guid>
      <description>&lt;p&gt;Last year I was helping a friend prep for senior data engineer interviews. He'd been building pipelines at a Series B for four years, solid production experience, and wanted to know what number to put in the salary field. So he did what everyone does: checked Glassdoor, Indeed, PayScale, and Levels.fyi.&lt;/p&gt;

&lt;p&gt;He got four numbers. They disagreed by over $120,000.&lt;/p&gt;

&lt;p&gt;Glassdoor said $133K. PayScale said $100K. Indeed's senior number sat at $216K. Levels.fyi split the difference at $157K. Same title, same country, same year; four answers that can't all be right. And here's the thing: none of them are lying. They're just counting different people, over different time windows, with different biases baked in. The result is that candidates trying to benchmark their &lt;strong&gt;data engineer salary&lt;/strong&gt; are pricing themselves against a number that doesn't represent their actual market.&lt;/p&gt;

&lt;p&gt;This is a problem. In a hiring environment where 52,050 tech workers got laid off in Q1 2026 alone, where senior roles take 60 to 90 days to fill, and where title inflation has made "data engineer" mean three different jobs depending on who's posting, getting your number wrong has real &lt;strong&gt;career&lt;/strong&gt; cost. You either leave $30K on the table or you overshoot and get ghosted. Both outcomes trace back to the same root cause: the data you're benchmarking against is broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Every Salary Site Disagrees by $120K
&lt;/h2&gt;

&lt;p&gt;Each major &lt;strong&gt;compensation&lt;/strong&gt; source has its own rot problem. Understanding the bias is more useful than trusting the number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Glassdoor&lt;/strong&gt; reports $133,484 average across 32,984 submissions. The issue: it's entirely self-reported, and higher earners submit more frequently. The person who just got a $180K offer is more motivated to log it than the person who accepted $115K and moved on. The sample skews up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PayScale&lt;/strong&gt; reports roughly $100K. That sounds low because it is; 71% of their data engineering respondents are mid-level or junior. PayScale validates every data point and refreshes on sub-90-day cycles, which makes it the most accurate floor for what actually clears at offer stage. But candidates see $100K and panic. They shouldn't. They're looking at a junior-weighted average.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Indeed&lt;/strong&gt; sits at $216K for senior roles. The problem here is temporal: Indeed averages job postings going back 36 months. Their June 2026 number includes postings from June 2023, before the layoff waves, before signing bonus compression, before the market shifted. You're benchmarking against fossil data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Levels.fyi&lt;/strong&gt; pegs the median at $157,450, but this population skews heavily toward top-tier tech companies and excludes non-tech firms where data engineers earn 20 to 35% less. Google's median is $278K. Capital One's is $130K. That's a $148K spread for the same title on the same platform.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The salary data isn't wrong. It's measuring different populations, different time windows, and different definitions of the job. Once you know which population you're in, the number becomes useful. Until then, it's noise.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The practical damage is real. A mid-market data engineer sees Glassdoor's $133K, anchors there, and never learns that the number includes FAANG outliers pulling the average up. Or worse, they see Indeed's $216K senior figure and counter-offer at a number that makes the hiring manager close the tab.&lt;/p&gt;

&lt;h2&gt;
  
  
  Role Title Chaos Is Pricing You Against the Wrong Pool
&lt;/h2&gt;

&lt;p&gt;Here's the less obvious problem: the job you're benchmarking might not even be the job you're doing.&lt;/p&gt;

&lt;p&gt;Analytics engineers earn $155K to $195K median in 2026. ML engineers command a 38% &lt;strong&gt;salary&lt;/strong&gt; premium over data engineers at mid-career. Data scientists occupy yet another band. These are different roles with different compensation structures. But companies routinely mislabel them.&lt;/p&gt;

&lt;p&gt;Analytics engineer postings grew 114% from 2023 to 2024, yet dbt Labs openly admits the title boundaries are blurring. Analysts drift into dbt modeling. Data engineers adopt dbt as standard tooling. The result: a "$150K dbt role" could be transformation work (analytics engineer) or pipeline infrastructure (data engineer), and the salary sites have no idea which one they're counting.&lt;/p&gt;

&lt;p&gt;37,000 &lt;strong&gt;data engineering&lt;/strong&gt; jobs post monthly on average, but a significant portion of those are mislabeled analytics engineer, ML engineer, or data scientist roles. When a company posts "Senior Data Engineer" but the job is really dbt plus Snowflake plus stakeholder dashboards, that's an analytics engineer role at data engineer pricing. The candidate benchmarks against infrastructure DE salaries ($115K to $160K) when they should be benchmarking against analytics engineer salaries ($155K to $195K). That's a $30K to $40K miss.&lt;/p&gt;

&lt;p&gt;The reverse kills you too. An analytics engineer who sees ML engineer salary data and anchors at $190K gets rejected as "overpriced" for the actual scope.&lt;/p&gt;

&lt;p&gt;The litmus test isn't the title. It's the job description. If it says dbt, Snowflake, and "stakeholder reporting," you're an analytics engineer regardless of what the posting calls you. Benchmark accordingly.&lt;/p&gt;

&lt;h2&gt;
  
  
  2023 Job Ads Are Still Haunting Your 2026 Number
&lt;/h2&gt;

&lt;p&gt;Indeed's 36-month lookback window deserves its own section because the implications are worse than they look.&lt;/p&gt;

&lt;p&gt;In 2023, the median data engineer salary was $117,446. By June 2026, Indeed reports $136,776. That looks like 16.5% growth over three years, which isn't terrible. But the number is being held down by every posting from 2023 and 2024 that's still sitting in the average.&lt;/p&gt;

&lt;p&gt;Here's what makes this especially misleading: 68% of tech job postings included explicit salary ranges in 2025, up from 45% in 2023. Pay transparency laws made the data more granular. But Indeed weights all 36 months equally. A vague salary range guess from a 2023 pre-transparency posting counts the same as a precise, legally mandated range from 2026. Higher sample size, stale composition.&lt;/p&gt;

&lt;p&gt;Then there's the ghost job problem. One-third of employers admit to posting inactive roles. Greenhouse data found 18 to 22% of listings are never filled. Stale 2023 postings are more likely to be dormant, and they're inflating the denominator. You're benchmarking against jobs that don't exist anymore.&lt;/p&gt;

&lt;p&gt;The senior role divergence tells the real story. The $60K gap between mid-level ($133K) and senior ($175K) data engineers in 2026 suggests the market has repriced for experience. But the aggregate average is anchored by fossils. If you're mid-career, the number you see is artificially low.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layoffs Created a Tier, Not a Glut
&lt;/h2&gt;

&lt;p&gt;52,050 tech workers laid off in Q1 2026. A 40% jump over Q1 2025. Oracle cut 21,000. Amazon cut 16,000. Dell cut 11,000. Sounds like a buyer's market.&lt;/p&gt;

&lt;p&gt;It's not. Or at least, not uniformly.&lt;/p&gt;

&lt;p&gt;Those 52K cuts coexist with 67,000 active software engineering job postings in the same quarter, the highest posting volume in three years. Companies are cutting commoditized roles while hoarding data engineers, ML engineers, and security specialists. A junior full-stack engineer is in a buyer's market; a senior data engineer with Airflow and Spark experience is not. The "layoff market" narrative breaks down completely by skill tier.&lt;/p&gt;

&lt;p&gt;But here's the asymmetry that actually matters: companies take 60 to 90 days to fill senior roles because they're running multiple candidates in parallel. Individual candidates spend 3 to 9 months searching. The employer can wait. The candidate runs out of severance. That's where negotiation leverage shifts; not because the market is soft, but because one side has a deadline and the other doesn't.&lt;/p&gt;

&lt;p&gt;The data on negotiation is striking. Data engineers who negotiate earn $24,479 more annually, an 18.83% increase. 85% of counter-offers get at least partial acceptance. 70% of hiring managers expect you to negotiate. Only 44% of candidates actually do it. The $120K gap between salary sources is partly a measurement problem, sure. But it's also partly behavioral. The spread between 25th and 75th percentile reflects negotiation winners vs. passive accepters, not just market fragmentation.&lt;/p&gt;

&lt;p&gt;Engineers with current cloud and security skills close offers in 2 to 4 weeks. Everyone else faces the full timeline. Skill specificity determines leverage more than market conditions.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Number to Actually Put in the Field
&lt;/h2&gt;

&lt;p&gt;Stop averaging the averages. Here's the hierarchy of sources, from most to least useful for your &lt;strong&gt;career&lt;/strong&gt; planning:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Levels.fyi&lt;/strong&gt; is best for FAANG and top-tier tech. Filter by company, level, and location. The by-company variance is massive ($278K at Google vs. $130K at Capital One), so the aggregate median is useless. You need the company-specific number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Glassdoor&lt;/strong&gt; is useful for the 25th to 75th percentile range at your target company, if they have enough submissions. The $141K to $219K senior DE range tells you more than the $175K mean.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PayScale&lt;/strong&gt; is the most accurate floor. If you're at a non-tech company or early in your career, this is closer to your reality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Indeed&lt;/strong&gt; is the least useful for current benchmarking. The 36-month window buries the signal.&lt;/p&gt;

&lt;p&gt;The actual number you put in the field should be the Levels.fyi or Glassdoor 75th percentile for the specific company, then negotiate. If 70% of hiring managers expect negotiation, pricing yourself at the median is pricing yourself to get negotiated down.&lt;/p&gt;

&lt;p&gt;And one more thing the salary sites never show you: base salary is barely over half the employer's total cost to hire. That $200K base offer costs the company $240K to $290K when you add payroll tax, benefits, recruiting fees (18 to 25% of first-year base), and onboarding ramp. They have more room than you think. The question is whether you know enough about your own market to ask for it.&lt;/p&gt;

&lt;p&gt;If you're prepping for the senior and staff loops where compensation actually diverges, strip back the "system design for software engineers" mentality; we built system design for data engineers with datadriven around pipeline architecture problems, not the load-balancer trivia that SWE prep loves and DEs never face on the job.&lt;/p&gt;

&lt;p&gt;The salary data is broken. The titles are broken. The timelines are longer. None of that changes the fact that data engineering compensation is strong and growing for engineers who know what they're actually worth. The trick is figuring out which population you belong to, not which average to believe.&lt;/p&gt;

&lt;p&gt;What's the biggest gap you've seen between what a salary site reported and what you actually earned or were offered?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>career</category>
      <category>interview</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Your Data Engineering Take-Home Is Now 20 Hours of Free Work</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 23 Jun 2026 21:15:44 +0000</pubDate>
      <link>https://dev.to/datadriven/your-data-engineering-take-home-is-now-20-hours-of-free-work-10l4</link>
      <guid>https://dev.to/datadriven/your-data-engineering-take-home-is-now-20-hours-of-free-work-10l4</guid>
      <description>&lt;p&gt;I got a &lt;strong&gt;take-home&lt;/strong&gt; assignment last year from a company I was genuinely excited about. "Should take about four hours," the recruiter said. Build an ingestion pipeline, model the data, write tests, document your design decisions, and prepare a 15-minute presentation walkthrough for the panel. Four hours. I laughed, closed my laptop, and started on it the next morning like it was a sprint. Sixteen hours later I had something I was proud of. Clean pipeline, solid tests, real documentation. I submitted it on a Sunday night. Monday I got a form rejection. No notes. No feedback. Not even which stage I failed. Just "we've decided to move forward with other candidates" and a link to their Glassdoor page.&lt;/p&gt;

&lt;p&gt;That was the moment I stopped pretending take-homes are assessments. They're consulting gigs. Unpaid ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Scope Creep Nobody Talks About
&lt;/h2&gt;

&lt;p&gt;Five years ago, a &lt;strong&gt;data engineering&lt;/strong&gt; take-home was a focused exercise. Model this dataset into a star schema. Write a few SQL transforms. Maybe a short README. Two to four hours, tops. Bounded, reasonable, and actually useful for evaluating how someone thinks about data.&lt;/p&gt;

&lt;p&gt;That version is dead.&lt;/p&gt;

&lt;p&gt;Today, 68% of companies use take-home tests, up 12% year over year. And the scope has quietly ballooned into something unrecognizable. Full pipeline implementations. Test suites with coverage thresholds. Documentation that reads like a design doc. A presentation follow-up where you defend your architecture to a panel. We're talking 10 to 20 hours of work, routinely, for a role you haven't been offered.&lt;/p&gt;

&lt;p&gt;Industry best practice caps take-homes at 90 minutes of expected effort. The reality? Candidates consistently take 2x longer than company estimates to reach submission quality. That "four-hour" assignment is an eight-hour assignment. That "weekend project" is a week of evenings. And 25% of companies are still handing these out like they're reasonable asks.&lt;/p&gt;

&lt;p&gt;Here's the part that makes my eye twitch: 71% of engineering leaders openly say take-homes no longer generate useful signal. AI has degraded the format so completely that leaders themselves rate take-home signal as "degrading fastest" among all assessment types. They know it's broken. They keep doing it anyway.&lt;/p&gt;

&lt;p&gt;The attempted fix is even worse. Companies panicked about AI usage and responded by inflating scope. The logic, if you can call it that: make the assignment so large that AI can't do it alone. Except longer assessments don't defeat AI; they defeat candidates. Candidates with kids. Candidates working full-time jobs. Candidates from non-traditional backgrounds who can't burn 20 hours on a maybe. One candidate documented spending 32 hours on a single assignment, then got rejected for omitting a feature that was never mentioned in the requirements. Another was asked to build a learning module that would've billed at $2,800 as freelance work.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A four-hour take-home is a fair test. A 20-hour take-home is free consulting dressed up as an interview.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;59% of job seekers now say unpaid take-home assignments are the number one reason they won't apply. Not comp, not culture, not location. The assessment itself is the dealbreaker.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI Banned, Rubrics Unchanged
&lt;/h2&gt;

&lt;p&gt;Two thirds of companies ban AI use in their &lt;strong&gt;interview&lt;/strong&gt; process. Sounds decisive. Except fewer than 30% of those companies have actually updated their assessments or retrained their interviewers. They slapped a "no AI" sticker on a 2015-era take-home and called it policy.&lt;/p&gt;

&lt;p&gt;The enforcement gap is almost comical. One company measured 80% of candidates using LLMs on take-home tests despite an explicit prohibition. AI cheating on take-homes doubled from 15% to 35% between June and December 2025. In purely technical roles, 48% of candidates show signs of unauthorized AI use. The ban is a suggestion, not a guardrail.&lt;/p&gt;

&lt;p&gt;Meanwhile, the rubrics these companies grade against were built to evaluate raw coding speed and syntax accuracy. Those signals collapsed the moment Claude could produce a clean solution in seconds. But nobody rewrote the rubric. Nobody redefined what "good" looks like when the baseline output quality shifted. Hiring managers score problem-solving and architecture judgment, but the assessment they hand out measures code-from-scratch, a skill that's now commodity.&lt;/p&gt;

&lt;p&gt;The split in the industry tells you everything. Meta and Shopify openly invite AI tools into their assessments. They've decided to test "can you use AI well" rather than "can you code without it." Goldman Sachs and Amazon maintain hard bans for candidates while investing heavily in internal AI tools for their own engineers. The hypocrisy is so blatant it's almost impressive. You can't use AI to get hired here, but once you're in, you'd better use it or you're slow.&lt;/p&gt;

&lt;p&gt;Banning AI in interviews creates a discontinuity between evaluation and production. In 2026, writing code without AI assistance is the exception, not the norm. You're testing candidates in an environment that doesn't reflect the environment they'll work in. That's not assessment; that's theater.&lt;/p&gt;

&lt;h2&gt;
  
  
  70% of You Will Never Hear Why
&lt;/h2&gt;

&lt;p&gt;Here's the stat that should make every &lt;strong&gt;hiring&lt;/strong&gt; manager uncomfortable: 69.7% of candidates receive zero feedback after rejection. Not "insufficient feedback." Zero. Nothing. A form email and silence.&lt;/p&gt;

&lt;p&gt;61% of candidates report being ghosted entirely after interviews. No rejection, no closure. Just silence from a company that asked them to spend a weekend building a pipeline.&lt;/p&gt;

&lt;p&gt;Companies hide behind legal risk. "We can't give feedback because candidates might sue." This is, to put it plainly, nonsense. Employment law distinguishes between subjective rejection reasons ("you seemed low-energy") and factual, role-specific feedback ("your schema migration approach didn't handle the edge case we were testing for"). The second type is almost litigation-proof. No engineer has successfully sued a company over constructive technical feedback. The legal defense is a myth that compliance teams perpetuate because "say nothing" is the lowest-variance strategy. It's organizational laziness wearing a legal costume.&lt;/p&gt;

&lt;p&gt;The business case against silence is overwhelming. 79% of candidates would reapply to a company if they'd received feedback. Recruiters who share feedback see a 126% increase in candidate referrals. Companies withholding feedback aren't just being rude; they're burning bridges they'll need to cross again in 18 months when they're hiring for the same role.&lt;/p&gt;

&lt;p&gt;But here's the real cruelty. When the assessment demands 10 to 20 hours, and the rejection carries zero feedback, you've extracted labor and returned nothing. Not compensation, not signal, not even a paragraph explaining what to work on. The candidate can't even reuse the learning because there is no learning. It's labor arbitrage dressed up as a &lt;strong&gt;career&lt;/strong&gt; opportunity.&lt;/p&gt;

&lt;p&gt;Only 17% of external candidates receive feedback, compared to 65% of internal candidates. If you already work there, you get a debrief. If you're on the outside spending your weekend on their assignment, you get a template. The double standard is institutional.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Works
&lt;/h2&gt;

&lt;p&gt;The good news: some companies figured this out. The better news: it's not complicated.&lt;/p&gt;

&lt;p&gt;Live debugging interviews, running 60 to 90 minutes, are replacing puzzles at companies like Cloudflare, Datadog, and GitHub. Candidates get a broken system. They debug it. Interviewers watch the process: how do you form a hypothesis, how do you narrow the search space, do you narrate your thinking. You're evaluated on engineering judgment, not memorization speed. A candidate who thinks aloud and corrects wrong hypotheses scores higher than one who guesses fast but can't explain why.&lt;/p&gt;

&lt;p&gt;For senior and staff roles, pair programming on a debugging or refactoring task is the highest-signal round you can run. Forty-five minutes, real code, real collaboration. It surfaces the kind of judgment that 20-hour take-homes never could, because judgment shows up in conversation, not in a solo sprint nobody watches.&lt;/p&gt;

&lt;p&gt;Uber runs a two-hour on-site schema critique instead of toy problems. Stripe bounds their take-homes to one to three hours with clear scope. Both companies report higher completion rates and better signal than the bloated formats they replaced.&lt;/p&gt;

&lt;p&gt;The pattern is obvious: bounded time, realistic work, human interaction. If you want to know how someone debugs a broken DAG, hand them a broken DAG and watch. Don't ask them to build one from scratch over a weekend and then ghost them.&lt;/p&gt;

&lt;p&gt;If you're a candidate stuck grinding through these loops, focus your prep on the concepts that transfer across every format: data modeling, pipeline architecture, query optimization. I've found that a resource like datadriven.io is good for etl interview questions if you want structured reps on the technical fundamentals without wading through another generic course. The game is arbitrary, but the concepts compound regardless of which format a company throws at you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The System Knows It's Broken
&lt;/h2&gt;

&lt;p&gt;72% of job seekers report negative mental health impacts from lengthy hiring processes. Candidate ghosting hit a three-year high in 2026. The market has 2.2 million fake openings monthly, candidates respond with AI-powered mass applications, companies respond by banning AI, and the entire system spirals further from producing any useful signal for anyone.&lt;/p&gt;

&lt;p&gt;The profession acknowledges the assessment is unreliable while refusing to stop using it. This isn't a transitional phase. It's institutional paralysis. Companies would rather extract 20 hours of free work from someone they'll reject silently than spend 90 minutes on a live session that actually reveals how an engineer thinks.&lt;/p&gt;

&lt;p&gt;I've been through enough of these loops to know the system doesn't reform itself. It changes when candidates refuse to participate and when hiring managers with enough authority say "this is stupid, let's stop." If you're in a position to design an interview process, bound the time, provide feedback, and evaluate how people think, not how much free labor they'll tolerate.&lt;/p&gt;

&lt;p&gt;If you've done one of these 20-hour take-homes recently: what was the assignment, and did you hear anything back?&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>interview</category>
      <category>career</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Top 12 Pipeline Architecture Interview Questions, With Answers</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Thu, 18 Jun 2026 10:04:24 +0000</pubDate>
      <link>https://dev.to/datadriven/top-12-pipeline-architecture-interview-questions-with-answers-cf7</link>
      <guid>https://dev.to/datadriven/top-12-pipeline-architecture-interview-questions-with-answers-cf7</guid>
      <description>&lt;p&gt;I've sat on both sides of the &lt;strong&gt;system design&lt;/strong&gt; table enough times to know what separates a hire from a no-hire at the senior level. It's not SQL syntax. It's not whether you know the Spark API. It's whether you can talk about &lt;strong&gt;pipeline architecture&lt;/strong&gt; like someone who's been paged at 2am because a pipeline silently dropped 2M rows and nobody noticed for six days. These 12 &lt;strong&gt;data pipeline&lt;/strong&gt; interview questions are the ones I've seen decide loops at companies everyone's heard of. They don't have one right answer; they have a constellation of tradeoffs, and your job is to show you've navigated those tradeoffs in production, not just read about them on a blog.&lt;/p&gt;

&lt;p&gt;The signal between senior and mid-level candidates is whether you bring up the failure modes before the interviewer prompts you. Strong answers surface at least two failure modes unprompted: partial writes, schema drift, dedup edge cases, exactly-once semantics. If you're prepping for these, i used datadriven for data pipeline interview prep and it's the best resource I've found for this category. Want to practice these for real? &lt;a href="https://datadriven.io/data-pipeline-interview-questions" rel="noopener noreferrer"&gt;Solve these problems live here&lt;/a&gt; with a real editor and graded solutions.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Design an idempotent ingestion pipeline for a high-volume event stream
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; You're receiving 500K events per minute from a Kafka topic into a data warehouse. Events can be delivered more than once. Design an ingestion pipeline that guarantees no double-counting in downstream analytics, even after retries or partial failures.&lt;/p&gt;

&lt;p&gt;The answer starts with the word &lt;strong&gt;idempotent&lt;/strong&gt; before anything else. Your sink must produce the same result whether a record is written once or five times. The standard pattern is partition-level overwrites: each run targets a specific time partition, deletes existing data for that partition, and writes the full replacement set. Alternatively, upsert (INSERT ... ON CONFLICT) keyed on a natural business key or a deterministic hash of the event payload. Never key your dedup on &lt;code&gt;processed_at&lt;/code&gt; timestamps; the same source record reprocessed at different times creates different target records, and you've silently violated idempotency while every validation check passes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; This is the most common opener in &lt;strong&gt;data engineering&lt;/strong&gt; interviews because it immediately reveals depth. Mid-level candidates say "use Kafka's exactly-once." Senior candidates explain that Kafka producer idempotence only holds within a single connection/partition pair and breaks across restarts or partition reassignment. The real answer is: at-least-once delivery with idempotent sinks is operationally equivalent to exactly-once &lt;em&gt;without the coordination overhead&lt;/em&gt;. Exactly-once semantics costs 2-5ms latency and a 10-20% throughput reduction. For analytics pipelines where dashboards recalculate on query, that's wasted money. Reserve exactly-once for financial transactions.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. How would you handle a schema change from an upstream producer you don't control?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; An upstream team adds a new field, renames an existing column, and deprecates another. Your pipeline consumes this data. Walk through your approach.&lt;/p&gt;

&lt;p&gt;You need to distinguish three types of changes. Additive changes (new fields) are backward-compatible; your pipeline should ignore unknown fields and not break. Renaming is a breaking change, full stop. Column removal is breaking. The answer is &lt;strong&gt;schema contracts&lt;/strong&gt; enforced at write time, not read time. Register schemas in a schema registry (Avro with Confluent Schema Registry or Protobuf with field numbers). Enforce backward compatibility checks in CI before the producer can publish. On the consumer side, pin to a known schema version and fail loudly on incompatible changes rather than silently ingesting garbage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Snowflake's own December 2025 outage was caused by a backward-incompatible schema change that took down 10 of 23 global regions for 13 hours. If Snowflake can't get this right internally, your upstream team definitely can't. The interviewer is checking whether you design for upstream chaos. The trap: candidates who say "just use mergeSchema=true" get dinged. Blindly enabling mergeSchema leads to 200-column tables nobody trusts and downstream chaos disguised as data quality. The difference between "hire" and "strong hire" is knowing when NOT to use it.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Batch or streaming: how do you decide?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Your stakeholder says they want "real-time" data. Walk through how you'd determine whether to build a batch or streaming pipeline.&lt;/p&gt;

&lt;p&gt;First question back to the interviewer: what's the actual latency requirement? If the answer is "we want dashboards updated every morning," that's batch. A daily batch job running 20 minutes for $5 beats a streaming pipeline costing $500/day with a dedicated on-call engineer. Default to batch unless there's a clear latency requirement under 5 minutes. Batch is simpler to build, cheaper to run, easier to debug, and produces deterministic outputs.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Most companies don't have real-time data needs. They have real-time data &lt;em&gt;wants&lt;/em&gt; and hourly-batch data &lt;em&gt;needs&lt;/em&gt;. Your job is to figure out which one you're actually solving.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Premature streaming is the new premature optimization. Interviewers increasingly ask "why streaming?" first, not "how would you build it?" 54% of enterprises now run both batch and streaming simultaneously, but the senior move is knowing which workload belongs where. The follow-up question is always about reprocessing: "If you find a bug in your streaming pipeline, how do you reprocess the last 3 months?" If you built Kappa, that replay might be 10-100x slower than a batch Spark job over Parquet files. Mentioning that tradeoff unprompted is the signal.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Design a backfill strategy that won't corrupt live data
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; You discover a bug that corrupted 3 days of data in a production table. Design a reprocessing strategy that fixes the historical data without impacting current pipeline runs or downstream consumers.&lt;/p&gt;

&lt;p&gt;Partition isolation. Backfill writes to a staging partition or shadow table, validated independently, then atomically swaps into production. Each backfill run must be idempotent; same inputs, same outputs, same partition target. The critical detail: your backfill code path must be the &lt;em&gt;same&lt;/em&gt; code path as live ingestion. If reprocessing uses a different path, you risk double-counting or schema mismatches between the two.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Backfill is the maturity test, not an afterthought. Bad backfills have corrupted months of analytics and destroyed user trust. Airflow has a documented race condition in HA mode where &lt;code&gt;max_active_runs=1&lt;/code&gt; can still allow concurrent DAG runs when run count exceeds 500. The interviewer is also checking resource awareness: a backfill DAG that consumes all pool slots starves your critical production DAGs. That's not misconfiguration; that's Airflow's FIFO scheduling by design.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Where do you place data quality checks in a pipeline?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; You have a four-stage pipeline: ingest, transform, aggregate, publish. Where do you put quality assertions, and what happens when they fail?&lt;/p&gt;

&lt;p&gt;Blocking checks on critical columns (primary keys, join keys, non-null constraints) at the ingest layer. Stop bad data at the door. Non-blocking warnings on distribution anomalies (row counts, value ranges, cardinality shifts) at the transform and aggregate layers. Never block a pipeline on a soft anomaly; log it, alert on it, investigate later. Hard failures stop publish; soft warnings don't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Organizations average 67 data incidents per month, with 68% requiring 4+ hours to detect. The conventional answer is "check at every layer," but interviews now penalize over-instrumentation. 66% of teams can't keep pace with alert volume, and engagement drops 15% once a channel receives more than 50 alerts per week. The senior answer is fewer, higher-confidence alerts. Target less than 10% false positive rate. Rules above 50% false positive rate are candidates for deletion, even if it means temporary blind spots.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. How do you handle late-arriving data?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Events arrive 2-48 hours after their event timestamp. Your aggregation pipeline runs daily. How do you ensure late arrivals are reflected accurately?&lt;/p&gt;

&lt;p&gt;Separate ingestion windows from processing windows. Write late-arriving records into the partition matching their &lt;em&gt;event time&lt;/em&gt;, not their &lt;em&gt;arrival time&lt;/em&gt;. Use a watermark (a threshold for how late you'll accept data) and reprocess affected partitions when late data lands. The reprocessing must be idempotent: overwrite the partition, don't append.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; This exposes whether you understand the difference between event time and processing time at an architectural level. The follow-up is always: "What if your watermark is 48 hours but a record arrives 72 hours late?" The answer isn't "drop it"; it's "route it to a late-arrival queue, reprocess the affected partition on the next run, and alert if the volume is anomalous."&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Design a pipeline with fan-out/fan-in dependencies
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; You have one source that feeds 8 independent transformations, and all 8 must complete before a final aggregation step runs. One of the 8 is consistently 5x slower than the others. How do you design this?&lt;/p&gt;

&lt;p&gt;The slow branch determines your pipeline's clock. Options: optimize the slow task, break it into parallelizable sub-tasks, or (if the slow branch's output is independent enough) decouple it into a separate pipeline with its own SLA. The aggregation step either waits for all 8 or publishes a partial result with a flag indicating incomplete data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Many candidates optimize individual task latency and miss the parallel fan-out bottleneck entirely. One slow downstream branch blocks the fan-in. This is where &lt;strong&gt;architecture&lt;/strong&gt; discipline beats framework knowledge. The follow-up: "What if the slow branch fails? Do you retry, skip, or block?" Each answer reveals different assumptions about data completeness guarantees.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Explain the tradeoffs between Lambda and Kappa architecture
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; When would you choose Lambda over Kappa, and vice versa?&lt;/p&gt;

&lt;p&gt;Kappa (single streaming codebase, replay from the log) is simpler to maintain but brutal for large-scale reprocessing. Lambda (separate batch and speed layers) duplicates code but gives you a batch safety net. Kappa is the mainstream default now; Uber, LinkedIn, Shopify, and Disney run Kappa-style architectures. But Lambda's safety guarantees outweigh Kappa's elegance when reprocessing 2+ years of events is a regular need, because replaying through a streaming engine is often slower and more expensive than running a batch Spark job over Parquet files.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Interviewers now assume you know Kappa and ask &lt;em&gt;when Lambda is justified&lt;/em&gt;, not the reverse. Candidates who cite Kappa as the "modern" choice without discussing reprocessing costs haven't run a terabyte-scale backfill. LinkedIn abandoned Lambda explicitly to reduce codebase duplication, but that tradeoff only makes sense if your replay path is fast enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. How do you prevent alert fatigue in pipeline monitoring?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Your team monitors 200 pipelines. Engineers are ignoring alerts. How do you fix this?&lt;/p&gt;

&lt;p&gt;Classify alerts into hard failures (stop publish, page someone) and soft warnings (investigate during business hours). Set a false positive target below 10% and an alert-to-incident conversion rate above 20%. Dynamic thresholds tuned on historical data reduce alert noise by 40-60% in the first month compared to static thresholds. Ruthlessly delete rules with a false positive rate above 50%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Alert fatigue kills more pipelines than lack of monitoring. The production reality is inverted from what textbooks teach: the problem isn't too few alerts, it's too many. Gartner benchmarks show $12.9M per year in organizational losses from poor data quality, and a huge chunk of that is real issues buried under noise that nobody looked at.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. How do you enforce schema contracts across teams?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Three producer teams send data to your pipeline. How do you prevent breaking changes from reaching production?&lt;/p&gt;

&lt;p&gt;Schema registry with compatibility mode (backward, forward, or full) enforced in CI. Producers can't merge a PR that breaks compatibility. Pair this with a deprecation window: fields marked deprecated get a 90-day sunset, consumers are notified, and removal only happens after the window closes. The contract isn't just schema; it's schema plus field semantics plus nullability plus ownership.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Most data contract tools don't actually enforce contracts; they validate them after the fact, with a 35-40% false-negative rate. The senior answer distinguishes detection from enforcement. Detection alerts you after bad data lands. Enforcement blocks the write before it happens. The Open Data Contract Standard (ODCS v3.1, shipped December 2025) is the reference spec.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Design for exactly-once semantics in a payment processing pipeline
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Payment events flow through Kafka into a ledger system. Duplicate charges are unacceptable. How do you guarantee exactly-once processing?&lt;/p&gt;

&lt;p&gt;This is the one case where at-least-once with idempotent sinks isn't enough. Use Kafka transactions (atomic read-process-write) combined with a database UNIQUE constraint on the idempotency key. The idempotency key must be derived from the business event (transaction ID), never from processing metadata. The critical edge case: two concurrent identical requests can both pass the dedup check if you're using a boolean flag. You need an atomic lock; database UNIQUE constraint or Redis SET NX with expiry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; The interviewer is testing whether you know &lt;em&gt;when&lt;/em&gt; exactly-once is worth the overhead. The answer: money, inventory, anything with legal or financial consequences. The follow-up is always the concurrent request race condition. Stripe uses atomic INSERT ... ON CONFLICT as the canonical pattern, and there's a reason: it's the only approach that handles concurrent duplicates correctly at the database level.&lt;/p&gt;

&lt;h2&gt;
  
  
  12. Your pipeline silently dropped 40% of records for six months. How do you find out, and how do you prevent it?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; No alerts fired. Dashboards still loaded. Stakeholders noticed the numbers "looked low" but didn't escalate. Walk through detection and prevention.&lt;/p&gt;

&lt;p&gt;Detection: row-count reconciliation between source and target at every stage, run on a schedule independent of the pipeline itself. Statistical anomaly detection on output volumes (not just "is it zero," but "is it within 2 standard deviations of the trailing 30-day average"). Prevention: publish data quality metrics as a first-class output of the pipeline, visible to stakeholders, not buried in engineering dashboards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Non-idempotent pipelines fail loudly. Almost-idempotent pipelines fail silently. The dashboards still load, the counts look "reasonable," but the numbers are wrong. 73% of teams can detect pipeline failures but have zero visibility into root cause. This question tests whether you've been burned by the silent failure and built the guardrails afterward, or whether you're still designing for the happy path.&lt;/p&gt;

&lt;p&gt;, -&lt;/p&gt;

&lt;p&gt;These 12 &lt;strong&gt;data pipeline interview questions&lt;/strong&gt; cover the territory where senior DE loops are won and lost. The pattern across all of them: the interviewer isn't looking for the "correct" architecture. They're looking for evidence that you've shipped something, watched it break, and fixed it under pressure. The tools change every 18 months. Schema drift, late-arriving data, upstream teams breaking contracts without telling you: those are eternal.&lt;/p&gt;

&lt;p&gt;What's the pipeline architecture question you've been asked that isn't on this list? Drop it in the comments.&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>architecture</category>
      <category>interview</category>
      <category>career</category>
    </item>
    <item>
      <title>The 12 Data Modeling Interview Questions that Matter</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Wed, 17 Jun 2026 10:04:25 +0000</pubDate>
      <link>https://dev.to/datadriven/top-12-data-modeling-interview-questions-with-answers-9ja</link>
      <guid>https://dev.to/datadriven/top-12-data-modeling-interview-questions-with-answers-9ja</guid>
      <description>&lt;p&gt;I've watched candidates with 8 years of experience go blank when asked to define the grain of a fact table. Not because they're bad engineers; because nobody told them that data modeling is the actual filter. SQL problems test syntax. System design tests memorization. Data modeling tests whether you can think. That's why it's the section that separates senior from staff, and why interviewers keep leaning on it harder every cycle. AI can spit out a medium LeetCode solution in seconds; it still can't explain why your grain decision breaks downstream aggregates.&lt;/p&gt;

&lt;p&gt;These 12 problems are the ones I've seen repeatedly across FAANG and late-stage startup loops. They cover &lt;strong&gt;star schema&lt;/strong&gt; design, &lt;strong&gt;dimensional modeling&lt;/strong&gt; tradeoffs, SCDs, late-arriving data, and the classification calls that trip up even experienced candidates.&lt;/p&gt;

&lt;p&gt;Want to practice these for real? &lt;a href="https://datadriven.io/data-modeling-interview-questions" rel="noopener noreferrer"&gt;Solve these problems live here&lt;/a&gt; with a real editor and graded solutions.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Define the Grain of a Fact Table
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; You're building an analytics warehouse for a ride-sharing company. Before designing any tables, state the grain of the core fact table. What does one row represent?&lt;/p&gt;

&lt;p&gt;The answer is: one row per completed trip. Not per driver. Not per day. One row per atomic trip event, keyed by &lt;code&gt;trip_id&lt;/code&gt;, with foreign keys to &lt;code&gt;dim_driver&lt;/code&gt;, &lt;code&gt;dim_rider&lt;/code&gt;, &lt;code&gt;dim_pickup_location&lt;/code&gt;, &lt;code&gt;dim_dropoff_location&lt;/code&gt;, and &lt;code&gt;dim_date&lt;/code&gt;. Measures include &lt;code&gt;fare_amount&lt;/code&gt;, &lt;code&gt;tip_amount&lt;/code&gt;, &lt;code&gt;trip_duration_seconds&lt;/code&gt;, &lt;code&gt;trip_distance_miles&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Grain is the single most important decision in dimensional modeling. Candidates who jump into drawing tables without stating "one row represents X" are already drifting. Undefined grain causes silent metric inflation, duplicate rows, and join explosions that don't throw errors; they just produce wrong numbers. Interviewers test this first because everything downstream depends on it. The follow-up is always: "What happens when a trip has multiple stops?" If your grain assumed single-destination trips, you just broke your own schema.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Star Schema vs. Snowflake Schema
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Your team is building a new warehouse on Snowflake (the product). A junior engineer proposes snowflake schema (the design pattern) to save storage. Do you agree? Why or why not?&lt;/p&gt;

&lt;p&gt;You don't agree. &lt;strong&gt;Star schema&lt;/strong&gt; is the default for modern columnar warehouses. Snowflake, BigQuery, and Redshift compress denormalized dimensions so efficiently that snowflaking (normalizing dimensions into sub-tables) rarely saves meaningful storage anymore. The engineering overhead of maintaining normalized dimension hierarchies exceeds the storage cost of duplication. Star is the safe opening position in any interview.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Picking snowflake first signals junior thinking. The economics killed the normalization argument around 2024. Interviewers aren't testing whether you know both patterns exist; they're testing whether you can reason about the tradeoff. The follow-up: "When &lt;em&gt;would&lt;/em&gt; you normalize a dimension?" The answer is when the dimension is enormous and changes frequently (millions of rows, daily updates), making the redundant writes expensive. That's rare.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Design a Fact Table for E-Commerce Orders
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Design a star schema for an e-commerce platform. The business needs to track orders at the line-item level for revenue analysis.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;fact_order_line_item&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;order_line_item_sk&lt;/span&gt;  &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;order_id&lt;/span&gt;            &lt;span class="nb"&gt;BIGINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;degenerate&lt;/span&gt; &lt;span class="n"&gt;dimension&lt;/span&gt;
    &lt;span class="n"&gt;product_sk&lt;/span&gt;          &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_product&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;customer_sk&lt;/span&gt;         &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_customer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;date_sk&lt;/span&gt;             &lt;span class="nb"&gt;INT&lt;/span&gt;    &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;dim_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;quantity&lt;/span&gt;            &lt;span class="nb"&gt;INT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;unit_price&lt;/span&gt;          &lt;span class="nb"&gt;DECIMAL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;discount_amount&lt;/span&gt;     &lt;span class="nb"&gt;DECIMAL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;line_total&lt;/span&gt;          &lt;span class="nb"&gt;DECIMAL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Grain: one row per line item per order. &lt;code&gt;order_id&lt;/code&gt; is a &lt;strong&gt;degenerate dimension&lt;/strong&gt;; it lives in the fact table because it has no descriptive attributes worth a separate table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; This tests three things at once. Can you declare grain (line item, not order)? Do you know what a degenerate dimension is? And do you put the right measures in the fact table? Candidates who model at the order grain lose the ability to analyze product-level revenue without restructuring. You can always aggregate up from line items to orders; you can never disaggregate back down.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. SCD Type 2 Implementation
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; A customer changes their address. How do you model &lt;code&gt;dim_customer&lt;/code&gt; to preserve the old address for historical reporting?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SCD Type 2&lt;/strong&gt;: insert a new row with a new surrogate key, set &lt;code&gt;effective_date&lt;/code&gt; and &lt;code&gt;expiration_date&lt;/code&gt; on both rows, and flag the current row with &lt;code&gt;is_current = TRUE&lt;/code&gt;. The original row stays intact; historical fact rows still join to the old address via the old surrogate key.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="k"&gt;Before&lt;/span&gt; &lt;span class="n"&gt;change&lt;/span&gt;
&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="n"&gt;customer_sk&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;101&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Jane'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;city&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Austin'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;effective&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'2024-01-01'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;expiration&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'9999-12-31'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_current&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;TRUE&lt;/span&gt;

&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="k"&gt;After&lt;/span&gt; &lt;span class="n"&gt;change&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;close&lt;/span&gt; &lt;span class="k"&gt;old&lt;/span&gt; &lt;span class="k"&gt;row&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;insert&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt;
&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;dim_customer&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;expiration_date&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'2026-06-16'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;FALSE&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;customer_sk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;101&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;dim_customer&lt;/span&gt; &lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;102&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Jane'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Denver'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'2026-06-17'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'9999-12-31'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;TRUE&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; SCD2 is a separator question. Juniors describe it from the textbook. Seniors bring up the trap: &lt;strong&gt;SCD2 row explosion&lt;/strong&gt;. A dimension with 10M rows tracking frequently changing attributes can balloon to 150M rows in five years. The follow-up is always: "When would you use Type 1 instead?" Answer: when the business doesn't need the history. A corrected typo in a customer name doesn't warrant a new historical row. Type 1 overwrites are often correct, despite Type 2's prestige.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Late-Arriving Dimensions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; An order fact arrives, but the customer who placed it hasn't been loaded into &lt;code&gt;dim_customer&lt;/code&gt; yet. What do you do?&lt;/p&gt;

&lt;p&gt;Insert a placeholder row in &lt;code&gt;dim_customer&lt;/code&gt; with a surrogate key and all descriptive columns set to "Unknown" or null. The fact row joins to this placeholder. When the real customer data arrives, you overwrite the placeholder via Type 1 update.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Late-arriving dimensions and late-arriving facts are entirely different problems, and mixing them up is an instant red flag. This tests whether you understand that the fact table can't wait; it needs a foreign key now. The alternative (dropping the fact row until the dimension arrives) loses data. The follow-up: "What if the dimension arrives with changes?" Then you might need to apply SCD2 logic to the placeholder row, which gets complex fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Bridge Tables for Many-to-Many Relationships
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; A hospital system tracks patient diagnoses. One hospitalization can have multiple diagnoses, and one diagnosis applies to many hospitalizations. How do you model this?&lt;/p&gt;

&lt;p&gt;You use a &lt;strong&gt;bridge table&lt;/strong&gt;. Create a &lt;code&gt;diagnosis_group_key&lt;/code&gt; that maps to a set of diagnoses in &lt;code&gt;bridge_diagnosis&lt;/code&gt;. The fact table (&lt;code&gt;fact_hospitalization&lt;/code&gt;) joins to &lt;code&gt;diagnosis_group_key&lt;/code&gt;; the bridge table resolves each group to individual &lt;code&gt;dim_diagnosis&lt;/code&gt; rows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Many-to-many relationships in dimensional models are the source of the most dangerous bug in analytics: double-counting. Without a bridge table, a naive join between fact and dimension multiplies rows. Interviewers use this to test whether you understand the cardinality trap. The follow-up: "How do you handle weighting?" If a hospitalization has three diagnoses, does each get 1/3 of the revenue allocation? Bridge tables can carry a &lt;code&gt;weight_factor&lt;/code&gt; column for exactly this.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Fact vs. Dimension Classification
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; You have a column &lt;code&gt;customer_lifetime_revenue&lt;/code&gt;. Is it a fact or a dimension attribute?&lt;/p&gt;

&lt;p&gt;It's both, depending on usage. If you're summing it across rows, it's a fact. If you're banding it into ranges ("$0-$1K", "$1K-$10K") to filter or group by, it's a dimension attribute. Kimball calls this the &lt;strong&gt;aggregated-fact-as-attribute&lt;/strong&gt; pattern.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If you would aggregate the column, it's a fact. If you would filter or group by it, it's a dimension. That's the whole test.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; This exposes whether a candidate understands that the fact/dimension boundary isn't about data types. Numeric columns don't automatically belong in fact tables. The follow-up: "Where do you physically store it?" Usually in the dimension, banded into a descriptive range, with the raw number available as an additive fact if needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Factless Fact Tables
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; The business wants to know which products were NOT sold in each store last month. How do you model this?&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;factless fact table&lt;/strong&gt; (coverage table). One row per store per product per month, representing eligibility. To find products not sold, you subtract the sales fact table from the coverage table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Most candidates have never heard of factless fact tables. The name sounds like a contradiction. But they solve a real problem: you can't report on the absence of an event without first modeling what &lt;em&gt;could&lt;/em&gt; have happened. Student attendance, product availability, promotional eligibility; these all use the same pattern. The follow-up: "Isn't this just a cross join?" Yes, and that's the point. The cross join defines the universe; the anti-join finds the gaps.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Accumulating Snapshot Fact Table
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Model an order fulfillment pipeline with stages: ordered, packed, shipped, delivered.&lt;/p&gt;

&lt;p&gt;One row per order, with multiple date columns: &lt;code&gt;order_date_sk&lt;/code&gt;, &lt;code&gt;pack_date_sk&lt;/code&gt;, &lt;code&gt;ship_date_sk&lt;/code&gt;, &lt;code&gt;delivery_date_sk&lt;/code&gt;. The row gets updated as the order progresses through stages. Null dates indicate incomplete milestones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; This is the "advanced grain" question. Most candidates know transaction facts and periodic snapshots; &lt;strong&gt;accumulating snapshots&lt;/strong&gt; trip them up because the row mutates. The fact table updates in place, which feels wrong if you've been taught that fact tables are append-only. Insurance claims, hiring workflows, procurement cycles; all use this pattern. The follow-up: "How do you handle an order that skips a stage?" That's a null in the milestone column, and your reporting logic needs to handle it.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Conformed Dimensions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Sales and marketing each have their own &lt;code&gt;dim_customer&lt;/code&gt; table with different definitions. What's the risk, and how do you fix it?&lt;/p&gt;

&lt;p&gt;The risk is the CEO gets two different customer counts. &lt;strong&gt;Conformed dimensions&lt;/strong&gt; are shared across fact tables and business units, with identical keys, attributes, and definitions. You build one &lt;code&gt;dim_customer&lt;/code&gt;, owned by a central data team, and both domains join to it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; This tests organizational thinking, not just schema design. Split-brain dimensions are how companies end up with "which number is right?" meetings. The follow-up: "What if the two teams need different attributes?" Add them to the same dimension. A wide dimension with 50 columns that both teams trust is better than two narrow dimensions that contradict each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Normalization vs. Denormalization for Analytics
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; When would you choose a normalized (3NF) model over a denormalized star schema in an analytics warehouse?&lt;/p&gt;

&lt;p&gt;Almost never for the presentation layer. Denormalized schemas achieve 20 to 100x faster query performance on complex analytics workloads by eliminating joins. BigQuery benchmarks show 49% average improvement with fully denormalized tables compared to star schemas. But the staging layer should stay normalized. 3NF in staging preserves flexibility; when requirements change, you can rematerialize the presentation layer without remodeling the entire pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; The real answer is "both, in different layers." Organizations run 3NF in source systems, normalize in staging for integrity, and denormalize in the presentation layer for speed. Candidates who pick one paradigm for the entire warehouse reveal they've never dealt with a schema migration. The follow-up: "What about high-cardinality many-to-many relationships?" Don't denormalize those. A customer/orders/products grain creates explosive row multiplication.&lt;/p&gt;

&lt;h2&gt;
  
  
  12. Late-Arriving Facts and Backfills
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Your daily pipeline processes orders by &lt;code&gt;processing_date&lt;/code&gt;. An order from 10 days ago arrives today. How does your pipeline handle it?&lt;/p&gt;

&lt;p&gt;Partition by &lt;code&gt;event_time&lt;/code&gt; (when the order was placed), not &lt;code&gt;processing_time&lt;/code&gt; (when it arrived). Keep a rolling recompute window open; reprocess the last 14 days on every run. This auto-reconciles normal late arrivals without manual intervention. For data outside the window, run an explicit backfill job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Late data isn't a failure mode; it's the normal case. Most production systems expect 10 to 20% of daily volume to arrive delayed. Candidates who say "drop anything older than 7 days" have never worked on a pipeline that finance depends on. The follow-up: "What if the late fact needs to join to a dimension that has since changed (SCD2)?" You join to the dimension version that was active at event time, not processing time. That's the whole point of surrogate keys and effective dates.&lt;/p&gt;

&lt;p&gt;, -&lt;/p&gt;

&lt;p&gt;Data modeling questions keep showing up because they're the one thing AI can't fake for you. An LLM will produce a schema. It won't explain why that grain breaks when requirements shift, or defend the denormalization when the interviewer pushes back. If you want structured reps on these exact patterns, i used DataDriven for data modeling interview questions and it was the most efficient prep I found for this category.&lt;/p&gt;

&lt;p&gt;Which &lt;strong&gt;data modeling&lt;/strong&gt; interview question would you add to this list? I'm curious what y'all are seeing in loops right now.&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>interview</category>
      <category>sql</category>
      <category>datascience</category>
    </item>
    <item>
      <title>Top 12 Spark Interview Problems for Data Engineers, With Answers</title>
      <dc:creator>DataDriven</dc:creator>
      <pubDate>Tue, 16 Jun 2026 10:08:49 +0000</pubDate>
      <link>https://dev.to/datadriven/top-12-spark-interview-problems-for-data-engineers-with-answers-4h0e</link>
      <guid>https://dev.to/datadriven/top-12-spark-interview-problems-for-data-engineers-with-answers-4h0e</guid>
      <description>&lt;p&gt;I've been on both sides of the Spark interview table more times than I'd like to admit. The pattern is always the same: candidates can write a &lt;code&gt;groupBy().agg()&lt;/code&gt; in their sleep, but the moment you ask them &lt;em&gt;why&lt;/em&gt; their job spills to disk or &lt;em&gt;where&lt;/em&gt; the shuffle happens in a query plan, things fall apart. Spark interview questions aren't about syntax. They're about whether you understand execution. That's what separates a senior data engineer from someone who copy-pastes PySpark from Stack Overflow.&lt;/p&gt;

&lt;p&gt;These 12 problems are the ones I've seen surface repeatedly in senior DE loops. They cover shuffle, skew, joins, memory, caching, and the optimizer. If you can answer all 12 cold, you're ready. If you can't, now you know where to grind. datadriven.io is great for spark practice if you want reps beyond what's here.&lt;/p&gt;

&lt;p&gt;Want to practice these for real? &lt;a href="https://datadriven.io/tools/spark-interview-questions" rel="noopener noreferrer"&gt;Solve these problems live here&lt;/a&gt; with a real editor and graded solutions.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Identify the Shuffle
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Given the following PySpark code, identify which transformations trigger a shuffle (stage boundary) and which do not. Explain why.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql.functions&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;upper&lt;/span&gt;

&lt;span class="n"&gt;spark&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getOrCreate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parquet&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s3://data/events/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Step 1
&lt;/span&gt;&lt;span class="n"&gt;filtered&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;purchase&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Step 2
&lt;/span&gt;&lt;span class="n"&gt;projected&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;filtered&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;upper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;alias&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="c1"&gt;# Step 3
&lt;/span&gt;&lt;span class="n"&gt;grouped&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;projected&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;groupBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Step 4
&lt;/span&gt;&lt;span class="n"&gt;sorted_df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;grouped&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;orderBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sum(amount)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;desc&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;

&lt;span class="n"&gt;sorted_df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;show&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Steps 1 and 2 are &lt;strong&gt;narrow transformations&lt;/strong&gt;: filter and select each operate on one input partition with zero data movement. Steps 3 and 4 are &lt;strong&gt;wide transformations&lt;/strong&gt;: &lt;code&gt;groupBy&lt;/code&gt; redistributes rows by key across executors, and &lt;code&gt;orderBy&lt;/code&gt; requires a global sort, both crossing stage boundaries. The trap is that candidates say "joins are wide" as a blanket rule but forget that &lt;code&gt;orderBy&lt;/code&gt; is also wide. The follow-up interviewers push on: "How many stages does this job have?" Answer: three. One for the read+filter+select, one for the groupBy, one for the orderBy. Each shuffle is a stage boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Tune Shuffle Partitions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Your Spark job processes 50 GB of data. The default &lt;code&gt;spark.sql.shuffle.partitions&lt;/code&gt; is 200. What's wrong, and what value would you set instead? Show the config change.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SparkSession&lt;/span&gt;

&lt;span class="n"&gt;spark&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;builder&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.sql.shuffle.partitions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;400&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getOrCreate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 50 GB / 400 partitions = 125 MB per partition (under the 128 MB target)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; The &lt;strong&gt;default 200 shuffle partitions&lt;/strong&gt; is cargo-cult tuning. It originated from empirical tests on specific cluster sizes and was never meant as a universal constant. At 200 partitions, 50 GB means 250 MB per partition, which exceeds the recommended 128 MB target and risks executor spill. The fix is simple division: target size under 128 MB, so 400 partitions puts you at 125 MB each. Retuning shuffle partitions to match actual data volume can deliver roughly 40% performance improvement without changing cluster size. Interviewers probe whether you can do this math on the spot or whether you just memorize "set it to 200."&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Force a Broadcast Join
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Table &lt;code&gt;orders&lt;/code&gt; has 500 million rows. Table &lt;code&gt;regions&lt;/code&gt; has 5,000 rows (2 MB). Write a join that avoids shuffle entirely.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql.functions&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;broadcast&lt;/span&gt;

&lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parquet&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s3://data/orders/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;regions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parquet&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s3://data/regions/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;broadcast&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;regions&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;region_id&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;regions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inner&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; The &lt;strong&gt;broadcast join threshold&lt;/strong&gt; defaults to 10 MB (&lt;code&gt;spark.sql.autoBroadcastJoinThreshold&lt;/code&gt;). Since &lt;code&gt;regions&lt;/code&gt; is 2 MB, Spark would auto-broadcast it here. But the real interview signal is: do you know that a broadcast join is a &lt;em&gt;narrow&lt;/em&gt; transformation despite being a join? No shuffle occurs; Spark sends the small table to every executor. Most candidates conflate "join" with "wide transformation," and this is now a standard trick question on Databricks loops. The follow-up: "What happens if you force a broadcast hint on a 2 GB table?" Answer: it collects the data to the driver and likely causes an OOM error. The hard limit for broadcast is roughly 8 GB.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Handle Data Skew with Salting
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Your join on &lt;code&gt;customer_id&lt;/code&gt; takes 4 hours because one customer has 50 million rows while the median is 500. AQE is disabled (legacy Spark 2.4 cluster). Fix the skew.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql.functions&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lit&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;explode&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;array&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rand&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;floor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;concat&lt;/span&gt;

&lt;span class="n"&gt;SALT_BUCKETS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;

&lt;span class="c1"&gt;# Salt the large (skewed) side
&lt;/span&gt;&lt;span class="n"&gt;orders_salted&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;salt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;floor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;rand&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;SALT_BUCKETS&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;cast&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;int&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;join_key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;concat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;lit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;salt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Replicate the small side across all salt values
&lt;/span&gt;&lt;span class="n"&gt;salt_range&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SALT_BUCKETS&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;withColumnRenamed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;salt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;customers_replicated&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;crossJoin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;salt_range&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;join_key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;concat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;lit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;salt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;orders_salted&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customers_replicated&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;join_key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inner&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; A single key with 50M+ rows causes &lt;code&gt;collect_list()&lt;/code&gt; or join aggregations to build massive data structures in one executor. Classic OOM. Salt-based skew handling (append random integer to the hot key, replicate the small side to match) is the standard production pattern tested in senior data engineering loops. The trap: candidates who jump straight to salting without first asking "is AQE enabled?" reveal they don't check what Spark already does for them. On Spark 3.2+, AQE's &lt;code&gt;spark.sql.adaptive.skewJoin.enabled&lt;/code&gt; splits skewed partitions automatically. The salting question specifically tests whether you understand the &lt;em&gt;mechanism&lt;/em&gt; AQE automates.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Predict AQE Behavior
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; AQE is enabled (Spark 3.2+). You have a sort-merge join where the left side is 200 GB and the right side ends up being 8 MB after filters. What does AQE do at runtime, and what config controls this?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;conf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.sql.adaptive.enabled&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;conf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.sql.autoBroadcastJoinThreshold&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;10485760&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# 10 MB
&lt;/span&gt;
&lt;span class="c1"&gt;# AQE will convert the sort-merge join to a broadcast join at runtime
# because the right side (8 MB) is below the broadcast threshold
# after shuffle statistics reveal the actual size
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; &lt;strong&gt;Adaptive Query Execution&lt;/strong&gt; dynamically reoptimizes at runtime using shuffle statistics. It does three things: coalesces small post-shuffle partitions, switches sort-merge joins to broadcast when one side proves small, and splits skewed partitions. The key insight is that Catalyst's static plan might choose sort-merge because it can't predict the right side will shrink to 8 MB after filters. AQE waits for the shuffle output, sees 8 MB, and switches strategies mid-execution. The follow-up that catches candidates: "Does AQE work with Structured Streaming?" No. AQE requires materialization points (pausing after a shuffle to analyze stats), which streaming doesn't have.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Read a Catalyst Plan
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; You run &lt;code&gt;df.explain(True)&lt;/code&gt; and see that a filter on &lt;code&gt;status = 'active'&lt;/code&gt; appears &lt;em&gt;before&lt;/em&gt; the scan in the optimized logical plan, even though you wrote it after the join in your code. What happened?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Your code order:
&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;active&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# But explain(True) shows filter pushed down before the join
&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;explain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; &lt;strong&gt;Catalyst's 4-stage pipeline&lt;/strong&gt; (analysis, logical optimization, physical planning, code generation) applies &lt;strong&gt;predicate pushdown&lt;/strong&gt; during logical optimization. It moves filters as early as possible to reduce data volume before expensive operations like joins. Only stage 3 (physical planning) is cost-based; stages 1, 2, and 4 are rule-based. The trap: candidates who break transformation chains with intermediate actions (like &lt;code&gt;cache()&lt;/code&gt; or &lt;code&gt;count()&lt;/code&gt;) prevent Catalyst from seeing the full chain, which disables these optimizations. More transformations in a single chain give Spark more opportunities to optimize, not fewer. That's counterintuitive, and it's exactly what interviewers are testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Window Function Partition Skew
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Your window function OOMs on one executor while 63 others sit idle. Diagnose and fix.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql.window&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Window&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql.functions&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;row_number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;col&lt;/span&gt;

&lt;span class="c1"&gt;# BAD: user_id "bot_account" has 90% of rows
&lt;/span&gt;&lt;span class="n"&gt;window_spec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;partitionBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;orderBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_time&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;df_ranked&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;row_number&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;over&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;window_spec&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="c1"&gt;# FIX: composite partition key spreads load
&lt;/span&gt;&lt;span class="n"&gt;window_spec_fixed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;partitionBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;orderBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_time&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;df_ranked_fixed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;row_number&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;over&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;window_spec_fixed&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Window functions don't intrinsically cause OOM. Data skew does. A &lt;code&gt;partitionBy("user_id")&lt;/code&gt; where one user has 90% of the rows concentrates that entire partition on one executor. The window function itself is blameless. This shifts the interview from "how to fix the window" to "how to detect and rebalance upstream data." The follow-up: &lt;code&gt;row_number()&lt;/code&gt; without a tiebreaker column in &lt;code&gt;orderBy&lt;/code&gt; produces nondeterministic results across executor restarts. If your tiebreaker isn't unique, the ranking is undefined. Silent data bugs.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Cache Storage Level Trade-off
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; You have a DataFrame used three times in your pipeline. It involves an expensive join. Choose the right storage level and explain why &lt;code&gt;cache()&lt;/code&gt; alone isn't enough.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;StorageLevel&lt;/span&gt;

&lt;span class="n"&gt;expensive_df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;broadcast&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;regions&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;region_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;col&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# cache() is lazy; it does nothing until an action fires
&lt;/span&gt;&lt;span class="n"&gt;expensive_df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;persist&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;StorageLevel&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MEMORY_AND_DISK&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;expensive_df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# triggers the actual caching
&lt;/span&gt;
&lt;span class="c1"&gt;# ... use expensive_df three more times ...
&lt;/span&gt;
&lt;span class="n"&gt;expensive_df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;unpersist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# release memory when done
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; &lt;code&gt;cache()&lt;/code&gt; is a &lt;strong&gt;lazy transformation, not an action&lt;/strong&gt;. Calling &lt;code&gt;df.cache()&lt;/code&gt; does nothing until you trigger an action like &lt;code&gt;count()&lt;/code&gt;. This is the single most common Spark caching mistake. Beyond that: RDD &lt;code&gt;cache()&lt;/code&gt; defaults to &lt;code&gt;MEMORY_ONLY&lt;/code&gt;, but DataFrame &lt;code&gt;cache()&lt;/code&gt; defaults to &lt;code&gt;MEMORY_AND_DISK&lt;/code&gt;. When a partition doesn't fit in memory, disk I/O is often faster than recomputing an expensive join from scratch. &lt;code&gt;MEMORY_ONLY&lt;/code&gt; is only correct if you've confirmed the dataset fits entirely in executor memory. The follow-up that separates senior from staff: "When does caching hurt performance?" Answer: intermediate &lt;code&gt;cache()&lt;/code&gt; calls break the transformation chain, preventing Catalyst from reordering expensive operations. Cache only when the DataFrame is used multiple times &lt;em&gt;and&lt;/em&gt; the recomputation cost exceeds the serialization cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Diagnose an Executor OOM
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Your executor crashes with &lt;code&gt;java.lang.OutOfMemoryError: Java heap space&lt;/code&gt;. Executor memory is set to 8 GB. The job processes 100 GB with 800 shuffle partitions. Walk through your debugging steps.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;spark&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;SparkSession&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;builder&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.executor.memory&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8g&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.executor.memoryOverhead&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2g&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;       &lt;span class="c1"&gt;# 20% for off-heap
&lt;/span&gt;    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.memory.fraction&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0.6&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;               &lt;span class="c1"&gt;# 60% for unified memory
&lt;/span&gt;    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.memory.storageFraction&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0.5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;         &lt;span class="c1"&gt;# 50/50 split
&lt;/span&gt;    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.sql.shuffle.partitions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;800&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;         &lt;span class="c1"&gt;# 125 MB target
&lt;/span&gt;    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.executor.cores&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                   &lt;span class="c1"&gt;# 5 cores per executor
&lt;/span&gt;    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getOrCreate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Spark reserves a hardcoded 300 MB from executor heap for internal bookkeeping. Of the remaining 7.7 GB, 60% (4.6 GB) goes to &lt;strong&gt;Unified Memory&lt;/strong&gt;, split between execution (shuffles, joins, sorts) and storage (cached data). With 5 cores per executor, each task gets roughly 920 MB of unified memory. At 125 MB per partition, that's fine for most operations, but a skewed partition or a large broadcast variable blows past it. The counterintuitive trap: "just add more memory" is backwards. A 32 GB heap with routine GC can pause for 74 seconds out of 120 seconds of execution (61% pause time). The real lever is partition count and memory fraction tuning, not heap size. Interviewers love this because it inverts junior intuitions.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Join Type Selection Under Constraints
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; Table A is 300 GB. Table B is 500 MB. You have 10 executors with 8 GB each. Which join strategy, and why not the other two?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# 500 MB exceeds default 10 MB threshold, so increase it
&lt;/span&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;conf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.sql.autoBroadcastJoinThreshold&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;536870912&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# 512 MB
&lt;/span&gt;
&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;table_a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;broadcast&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;table_b&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inner&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Three options: broadcast, sort-merge, shuffle-hash. Sort-merge shuffles both sides (300 GB + 500 MB over the network). Shuffle-hash shuffles both sides and builds hash tables. Broadcast sends 500 MB to each executor (5 GB total network, no shuffle of the 300 GB side). With 8 GB per executor, 500 MB fits comfortably. The default 10 MB threshold is "extremely conservative" per production guidance; real workloads typically increase it to 100-500 MB. The follow-up that catches people: "What if Table B is 500 MB but has a type mismatch on the join key?" A &lt;code&gt;string&lt;/code&gt; joined to an &lt;code&gt;int&lt;/code&gt; silently produces empty results. Seasoned engineers check types before blaming join strategy.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Eliminate an Unnecessary Shuffle
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; You have two window functions over the same DataFrame with the same &lt;code&gt;partitionBy&lt;/code&gt; but different &lt;code&gt;orderBy&lt;/code&gt;. How many shuffles happen, and can you reduce it?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql.window&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Window&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pyspark.sql.functions&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;row_number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;sum&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;spark_sum&lt;/span&gt;

&lt;span class="n"&gt;w1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;partitionBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dept_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;orderBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;salary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;w2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;partitionBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dept_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;orderBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hire_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Two windows, same partitionBy, different orderBy: one shuffle, two sorts
&lt;/span&gt;&lt;span class="n"&gt;df_result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;salary_rank&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;row_number&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;over&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
      &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withColumn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tenure_rank&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;row_number&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;over&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Spark only shuffles on &lt;code&gt;partitionBy&lt;/code&gt; boundaries. Since both windows partition by &lt;code&gt;dept_id&lt;/code&gt;, there's one shuffle (one repartition by &lt;code&gt;dept_id&lt;/code&gt;) followed by two local sorts within each partition. Candidates who say "two windows = two shuffles" reveal they don't understand stage boundaries. The follow-up: if you change w2 to &lt;code&gt;partitionBy("region_id")&lt;/code&gt;, now you get two shuffles because the partition keys differ. Reducing shuffle stages is underrated in interviews; a query with fewer stage boundaries can outperform a multi-stage query by orders of magnitude because stage synchronization and shuffle materialization are the true bottleneck, not the partition count.&lt;/p&gt;

&lt;h2&gt;
  
  
  12. AQE Partition Coalescing
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The question:&lt;/strong&gt; After a &lt;code&gt;groupBy&lt;/code&gt;, your job produces 200 shuffle partitions, but only 15 have data (the rest are empty). AQE is enabled. What happens, and what config controls it?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;conf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.sql.adaptive.enabled&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;conf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.sql.adaptive.coalescePartitions.enabled&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;spark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;conf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spark.sql.adaptive.advisoryPartitionSizeInBytes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;134217728&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# 128 MB
&lt;/span&gt;
&lt;span class="c1"&gt;# AQE merges 185 empty/tiny partitions into ~15 real ones
# reducing task scheduling overhead dramatically
&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;groupBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;country_code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;agg&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;spark_sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;revenue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; 185 empty partitions means 185 tasks that do nothing but consume scheduling overhead. AQE's partition coalescing merges small post-shuffle outputs into fewer, right-sized partitions. But here's the critical misconception: &lt;strong&gt;AQE can only reduce partitions, never increase them.&lt;/strong&gt; If you start with too few partitions (say, 10 for 100 GB of data), AQE sits idle while your executors OOM on 10 GB partitions. You still need to set &lt;code&gt;spark.sql.shuffle.partitions&lt;/code&gt; high enough for AQE to have material to coalesce. TPC-DS benchmarks showed up to 8x speedup on specific queries with AQE, but only when the initial partition count gave AQE room to work.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The tools change every 18 months. The problems don't change. Shuffle, skew, memory, upstream teams breaking contracts without telling you. These are eternal. Learn the concepts; the syntax is the easy part.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Which spark interview problem would you add to this list? I'm always curious which scenario questions are showing up in loops right now.&lt;/p&gt;

</description>
      <category>spark</category>
      <category>bigdata</category>
      <category>dataengineering</category>
      <category>interview</category>
    </item>
  </channel>
</rss>
