A recent r/dataengineering discussion started with a frustrating result. A company interviewed 22 people for a data engineering role and could not find a match. After working to create data teams (specifically data engineering teams) for over a decade, I feel like this reflected a lot of the observations I also have had over these years.
The job needed someone with software engineering fundamentals who could work on a lakehouse and integrate it with the company's main application. The requested stack included Spark, AWS Glue, Lambda, Airflow, and dbt. Some applicants had analytics backgrounds, some came from software engineering without Spark experience, and some qualified data engineers wanted more money than the budget allowed.
The comments quickly became a debate about salary, training, and what the title "data engineer" is supposed to mean. That debate is useful because the hiring problem was not really a shortage of keywords. The role combined several kinds of ownership without saying which ones mattered most.
If a team needs one person to model business data, build distributed processing jobs, write application integrations, manage cloud infrastructure, operate production pipelines, and know its exact toolchain, it may be trying to hire a small team through one job description.
I would fix the role before adding another interview round.
A tool list does not describe the job
Spark, Glue, Lambda, Airflow, and dbt can appear in the same platform. They do not represent one coherent skill.
Someone may be excellent at dimensional modelling and dbt but have limited experience writing a service that exposes data to an application. Another engineer may understand distributed systems, IAM, deployment, and failure recovery but need a few weeks to learn Glue. A third may know every service name in the description because their last company used the same stack, yet struggle to explain what happens when a pipeline partially writes its output and retries.
The third candidate often passes the keyword screen. The first two may be more useful once the system reaches production.
AWS's Data Analytics Lens is a good reminder of how broad production data work becomes. Its design principles cover operational excellence, security, reliability, performance, cost, and governance of data and metadata changes. Knowing the SDK call that starts a Glue job is a small part of that responsibility.
I would define the role through the system it owns:
Then I would mark the boundaries this engineer owns directly, the boundaries they influence, and the boundaries another team operates.
That makes the requirement much clearer. "Own ingestion through serving for the customer lakehouse" describes a job. "Must know Spark, Glue, Lambda, Airflow, dbt, Terraform, Kafka, and Snowflake" describes a shopping trip.
Software engineering fundamentals need a concrete meaning
"Clean code" and "DRY" sound reasonable in a job description. They are also difficult to assess because two experienced engineers can disagree about both.
For a data platform, I would translate software engineering fundamentals into observable behaviours:
- Can the candidate split a pipeline into components with clear inputs and outputs?
- Can they make a write idempotent or explain why it cannot be?
- Can they test transformation logic without deploying the whole platform?
- Can they evolve a schema without surprising downstream consumers?
- Can they trace a failed run, identify the damaged data, and describe recovery?
- Can they review a change for security, cost, and operational impact?
These questions are closer to the work. They also give candidates from adjacent backgrounds a fair chance to show transferable judgement.
Google's SRE guidance treats a data pipeline as a service. The Data Processing Pipelines chapter recommends defining customer-facing objectives, detecting freshness and correctness failures, documenting the system, and building a development lifecycle that catches problems before production. That is software engineering applied to data. It is not a particular folder structure or an argument about whether a helper function should contain six lines or eight.
The interview should therefore test a change over time, not a snapshot of remembered syntax.
Give the candidate a small pipeline design. Then change one assumption:
The useful signal is how the candidate updates the design. Do they notice that corrections change the merge strategy? Do they separate freshness from correctness? Do they ask about deletion propagation, backups, and derived datasets? Do they add streaming infrastructure immediately, or first test whether frequent incremental batches meet the requirement?
That conversation tells me more than asking someone to name five Spark transformations.
Decide what must be known on day one
Every role has skills that cannot wait. The mistake is treating the entire current stack as one of them.
I would divide requirements into three groups:
| Group | Meaning | Example |
|---|---|---|
| Required judgement | Hard to acquire during onboarding | Failure recovery, data correctness, security boundaries |
| Transferable implementation skill | Needed, but not tied to one vendor | Distributed processing, orchestration, SQL, testing |
| Local stack knowledge | Can be learned with access and support | Glue configuration, repository conventions, one dbt package |
This does not mean tools never matter. A team migrating a complex Spark estate may reasonably need someone who has debugged Spark executors under load. A three-month delivery deadline may leave little time for cloud onboarding. Those are constraints, and the job description should say so.
The list becomes suspicious when every item is mandatory but none is identified as decisive.
For a mid-level lakehouse role, I would usually make production pipeline ownership and solid programming ability mandatory. I would require depth in one distributed processing environment, not necessarily the exact managed service. I would treat the local orchestration and transformation tools as learnable unless the person will own them alone from their first week.
If the team cannot support any learning, that is another requirement worth stating. It also means the role is more senior than the number of years in the description may suggest.
A better interview loop
I would keep the loop small and connect every stage to a decision.
The first conversation should establish scope. Ask the candidate to describe a data system they owned, where their responsibility started, where it ended, and what failure they were expected to handle. This prevents a familiar product name from standing in for actual ownership.
The technical exercise should be a short design problem using a domain anyone can understand. Orders, invoices, or device readings are enough. Provide volumes, latency, and recovery requirements. Let the candidate ask for missing constraints.
The code review should contain an ordinary pipeline bug: a non-idempotent retry, a timestamp interpreted in local time, an unbounded read, or a write that can expose partial output. Ask what they would change and how they would prove the fix.
The operational discussion should begin with evidence:
The job is green.
The dashboard is missing yesterday's revenue.
What do you inspect first?
A good answer separates orchestration status from data correctness. It moves through source arrival, processing boundaries, output counts, freshness, and reconciliation. The exact monitoring product is secondary.
Finally, include one learning question. Pick a tool the candidate has not used and provide a short piece of documentation. Ask them to explain how they would validate it in a proof of concept. This tests the skill the job will require repeatedly: learning the next service without pretending its logo is an architecture.
The role may still be too wide
Improving the interview cannot repair an incoherent operating model.
If one engineer is expected to own business definitions, application APIs, Spark performance, IAM, Terraform, on-call response, and stakeholder reporting, the team needs to choose. It can narrow the role, split the responsibilities, raise the seniority and compensation, or hire for potential and provide support.
There is no interview technique that creates a candidate who is simultaneously specialised in every layer, available within the current budget, and productive without onboarding. Sometimes the market is giving feedback about the design of the job.
My recommendation is to hire for the hardest judgement the team cannot teach quickly. Make the owned system explicit. Test engineering through realistic changes and failures. Treat vendor knowledge as a preference unless there is a concrete reason it must exist on day one.
The goal is not to find someone who has seen the same collection of tools. It is to find someone who can keep the data system correct when the collection changes. That is a rational and practical approach to a never-ending cycle.


Top comments (0)