What the research found
A 2025 study in PNAS Nexus tested five widely used large language models: OpenAI’s GPT-3.5 Turbo and GPT-4o, Google’s Gemini 1.5 Flash, Anthropic’s Claude 3.5 Sonnet, and Meta’s Llama 3-70b, in a randomized experiment scoring about 361,000 entry-level job resumes with randomly assigned social identities attached. Across the board, the models awarded higher scores to female candidates with the same work experience, education, and skills as their male counterparts, while giving lower scores to Black male candidates with comparable qualifications to other groups. The researchers estimate this played out as hiring probability differences of about 1 to 3 percentage points for otherwise identical candidates, a pattern that held steady across different job positions and subsamples tested.
The direction and size of the bias varied model to model, but each model tested showed a measurable bias in one direction or another. None of the five scored candidates in a way that stayed neutral to identity once other qualifications were held equal.
Why holding qualifications equal is what makes this finding matter
Bias studies that compare resumes that differ in the real world always leave room for a defense: maybe the more favored group’s resumes were a bit stronger in some way the comparison missed. This study closes that door by design: the same experience, education, and skills were tested across different assigned identities, with only the name and demographic signal changed. Whatever difference in scoring showed up came from the identity attached to the resume, not the substance of it. That’s the part that makes this a bias finding rather than an ambiguous pattern with another possible explanation.
Why this matters even if you’re not a hiring manager
AI resume screening is in active use across hiring pipelines, filtering thousands of applications down to the ones a human recruiter sees. If you’ve applied for a job and gotten no response with no clear reason why, an automated screening step is a plausible part of that pipeline, and this research says that step isn’t neutral. It’s a documented property of systems that function close to how they’re already used in hiring today, not a hypothetical concern about some future use of AI.
The practical takeaway
If you’re on the hiring side, treat AI resume screening as a tool that needs active auditing for bias, not something to deploy and trust by default, and consider periodic checks using matched, identity-varied test resumes similar to how this study was designed. If you’re job hunting, this is one more reason a resume that reads as strong on paper doesn’t guarantee a fair first look, a frustrating reality but not one you can fix through how you write a resume alone.
This research tested specific model versions at a specific point in time. AI companies update and retrain their models often, and a model’s behavior on this exact test could shift with newer versions. The consistent pattern across five different companies’ systems, rather than any single model’s specific numbers, is the part of this finding most likely to hold up over time.
Reference
An, J., Huang, D., Lin, C., Tai, M. “Measuring gender and racial biases in large language models: Intersectional evidence from automated resume evaluation.” PNAS Nexus, 2025. https://doi.org/10.1093/pnasnexus/pgaf089

Top comments (0)