The strange thing wasn't that more candidates were passing.
The strange thing was that nobody thought it was strange.
Six months after we launched our AI interview system, everyone was happy.
Recruiters were happy.
Hiring managers were happy.
Candidates were happy.
Even the CTO was happy.
The numbers looked fantastic.
- Candidate pass rate: up 23%
- Interview completion rate: up 31%
- Candidate satisfaction: up 18%
Every dashboard was green.
Every weekly report looked better than the week before.
The AI interviewer seemed to be working exactly as intended.
Then someone asked an uncomfortable question.
If we're hiring better people, why aren't our teams getting stronger?
Nobody had a good answer.
The Pattern
The question came from a senior engineering manager.
Not from HR.
Not from leadership.
Not from the AI team.
He had reviewed the performance evaluations of engineers hired during the previous two quarters.
The data wasn't catastrophic.
But it wasn't improving either.
New hires were passing interviews more often.
Yet their first six months looked almost identical to previous cohorts.
The hiring funnel was getting easier.
The outcome wasn't getting better.
At first glance, it looked like normal variation.
Then we graphed the numbers.
The pass rate wasn't just increasing.
It was increasing steadily.
Month after month.
Almost like someone was turning a dial.
The Obvious Suspects
The first explanation was ChatGPT.
Maybe candidates had simply become better at using AI tools.
We checked.
No correlation.
The increase affected coding questions, system design questions, and behavioral questions equally.
That was odd.
Generative AI usually improves some categories more than others.
Not everything at once.
The second explanation was question leakage.
Maybe interview questions had spread online.
We rotated the question bank.
Nothing changed.
Pass rates kept rising.
The third explanation was model drift.
Perhaps the scoring model had gradually become more lenient.
The model version hadn't changed in months.
No retraining.
No parameter updates.
Nothing.
The system wasn't becoming easier because of the model.
Something else was moving.
Following the Scores
We stopped looking at candidates.
We started looking at interviewers.
Not human interviewers.
Human reviewers.
Our AI system generated recommendations.
Humans provided feedback.
Those feedback signals were later used to calibrate future scoring.
The design seemed reasonable.
The AI wasn't making final decisions.
Humans still approved or rejected candidates.
The system simply learned from those outcomes.
At least that was the theory.
One recruiter stood out.
Not because she reviewed more candidates.
Not because she hired more candidates.
Because her feedback disagreed with everyone else.
Consistently.
Most recruiters approved around 55–65% of AI recommendations.
She approved nearly 90%.
Every month.
Every quarter.
Every role.
Engineering.
Product.
Data.
Didn't matter.
Her approval rate barely moved.
At first we assumed she was simply optimistic.
Then we discovered something else.
The AI loved candidates she loved.
Far more than statistics would predict.
The Loop
The scoring model wasn't retraining itself directly.
That would have been easy to find.
Instead, every month we recalibrated scoring thresholds using historical hiring outcomes.
The logic seemed harmless.
If recruiters consistently approved certain profiles, the system adjusted future recommendations accordingly.
A feedback mechanism.
Designed to align AI behavior with human decisions.
The problem was scale.
One recruiter reviewed nearly three times as many candidates as everyone else.
Not because she was more senior.
Because she worked across multiple departments.
Over time her decisions became disproportionately represented in the calibration data.
Nobody noticed.
The dashboards tracked model accuracy.
They tracked candidate satisfaction.
They tracked hiring speed.
Nobody tracked influence concentration.
The AI wasn't learning from the company.
The AI was slowly learning from one person.
The Experiment
We replayed six months of interview data.
Then removed her feedback entirely.
The results surprised everyone.
The pass rate curve flattened almost immediately.
Several recommendation patterns disappeared.
The model's definition of a "strong candidate" shifted.
Not dramatically.
Just enough.
Enough to explain nearly all of the growth we had observed.
The system hadn't become smarter.
It had become more aligned with one recruiter's preferences.
And here's the uncomfortable part.
She wasn't wrong.
Many of her hires performed perfectly well.
Some became top performers.
The issue wasn't quality.
The issue was representation.
One person's judgment had become company policy.
Nobody voted on it.
Nobody approved it.
Nobody even knew it was happening.
The Hardest Bug to Explain
When we presented the findings, nobody argued with the data.
The graphs were clear.
The mechanism was clear.
The feedback loop was clear.
But one executive asked a question that stayed with me.
If her judgment was good, why is this a problem?
The room went quiet.
Because technically, it wasn't a bug.
The system worked exactly as designed.
Humans provided feedback.
The AI learned from feedback.
The AI became better at predicting future feedback.
Mission accomplished.
The problem was that we had confused consistency with objectivity.
The AI wasn't discovering what made a great engineer.
It was discovering what made one recruiter say yes.
Those are not the same thing.
What We Changed
We didn't remove human feedback.
That would have been the wrong lesson.
Instead we started measuring something we had never measured before.
Influence.
Not accuracy.
Not acceptance rate.
Influence.
Who shapes the training signal?
How concentrated is that influence?
How much of the system's future behavior can be traced back to a single source?
Those metrics never appeared on executive dashboards before.
Now they do.
The Real Lesson
The most dangerous AI failures rarely look like failures.
No outage.
No incident.
No security breach.
No alarming stack trace.
Everything was green.
Every KPI was improving.
Every report suggested success.
That's what made this one hard to see.
The system wasn't broken.
The system was becoming more confident in a perspective that nobody realized it had inherited.
The interviews weren't getting easier because candidates were improving.
They weren't getting easier because the model was drifting.
They weren't getting easier because somebody manipulated the system.
They were getting easier because one person's idea of talent had quietly become the AI's idea of talent.
And nobody noticed until the numbers looked too good to question.
Top comments (0)