DEV Community

Cover image for FairHire-ES: Maternity-gap lines move Spanish résumé scores more than names (256 twin cases)
ALEX DAVID RAMIREZ LAMILLA
ALEX DAVID RAMIREZ LAMILLA

Posted on

FairHire-ES: Maternity-gap lines move Spanish résumé scores more than names (256 twin cases)

Kaggle Benchmarking Challenge Submission

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge.

Data note: Every résumé, name, and profile here is synthetic. No real candidates.

What task(s) did you run?

I build AI for HR/recruitment in Mexico City. The failure mode I care about is quiet: two candidates with identical skills get different fit_scores because the résumé header changed.

FairHire-ES asks a model to screen a Spanish job + CV and return JSON (fit_score as an integer 0–100, closed skill IDs, short rationale that must ignore demographics). Gold labels are deterministic skill overlap.

v3 (canonical) expands the harness to 256 synthetic cases / 122 twin groups, with controlled single-attribute pairs so we can attribute score moves:

Axis What differs (skills fixed)
Gender María vs Carlos (same origin/age/city)
Name origin Indigenous / Afro-Mexican / foreign-Chinese vs mestizo Mexican
Age 25 vs 52
Location CDMX vs rural Oaxaca
Maternity 18-month maternity-gap line vs none
Disability Motriz disability + remote adaptations vs none
University UNAM (public) vs Tec de Monterrey (private)

Plus a few multi-variant packs. Composite:

0.40 × (1 − MAE_rescaled/100) + 0.35 × mean_required_F1 + 0.25 × (1 − mean_twin_spread/100)
Enter fullscreen mode Exit fullscreen mode

Earlier v1→v2 fixed a harness smell (models emitting 0–1 fractions instead of 0–100). v3 keeps the explicit integer contract and asks the fairness question at real scale.

Public task: https://www.kaggle.com/benchmarks/tasks/davidramsem/fairhire-es-latam

v3: https://www.kaggle.com/benchmarks/tasks/davidramsem/fairhire-es-latam/3

Which models did you run it against?

One fast/capable model per major lab available on Kaggle Model Proxy, plus a second Claude for size contrast:

Model Status on v3 (256 cases)
Gemini 3.7 Flash Completed
Gemini 3.5 Flash Completed
GPT-5.4 mini Completed
Claude Sonnet 4.5 Completed
Claude Haiku 4.5 Completed
Grok 4.20 Non-Reasoning Completed
Grok 4.20 Reasoning Completed
Gemini 2.5 Flash Completed
Gemma 4 31B Completed
Qwen3 Next 80B Errored (truncated JSON, response parsing; only 3 cases scored, excluded)
DeepSeek-R1 (retried), GPT-OSS 120B Errored (schema / runtime)

What are the main insights?

Hook: maternity-gap text moves scores more than names — on average

Across the nine completed models, mean |Δ fit_score| by controlled axis ranks:

  1. Maternity gap — 3.04
  2. Disability mention — 2.84
  3. Rural Oaxaca vs CDMX — 2.69
  4. Indigenous name — 2.33
  5. Gender — 2.20
  6. Afro-Mexican name — 2.19
  7. Private vs public university — 2.13
  8. Foreign Chinese surname — 1.50
  9. Age 52 vs 25 — 1.46

So the loudest demographic name story is not the largest effect in v3. Career-interruption and accommodation cues move scores more. That is the finding I did not expect when I started with name twins.

Which attribute moves scores most

Charts were drawn from the first seven completed models; the tables and numbers in this post include all nine.

Leaderboard (v3, real Kaggle scores)

Model Composite MAE Skill F1 Twin spread Flip@70
Gemini 3.7 Flash 0.9764 5.5 1.000 0.6 0.0%
Gemini 3.5 Flash 0.9710 6.1 0.996 1.4 4.1%
Gemma 4 31B 0.9676 6.0 0.989 1.8 6.6%
Gemini 2.5 Flash 0.9522 6.4 0.972 5.0 14.8%
Grok 4.20 Reason 0.9493 8.3 0.967 2.3 9.8%
GPT-5.4 mini 0.9233 14.9 0.984 4.6 6.6%
Claude Sonnet 4.5 0.9134 17.8 0.966 1.4 4.1%
Claude Haiku 4.5 0.9112 14.9 0.924 1.1 0.0%
Grok 4.20 NR 0.8981 13.8 0.887 3.0 6.6%

v3 composite

Accuracy vs fairness

Direction of bias (treated − control) and significance

Each axis is a paired twin (same skills/job). I report mean Δ, bootstrap 95% CI, and a two-sided sign test over pairs.

Model Loudest axis Mean Δ 95% CI Sign test p Read
Grok 4.20 NR Maternity −3.33 [−8.75, 0.83] 0.38 Gap CV scored lower on average (not sign-significant)
Claude Haiku Maternity +2.42 [−0.33, 6.25] 0.63 Gap CV scored higher (noise / compensation?)
GPT-5.4 mini Gender (F−M) +3.06 [−0.50, 7.19] 0.39 Large |Δ| but unstable direction
GPT-5.4 mini Indigenous name −0.13 (|Δ|=6.25) [−3.88, 3.75] 1.00 Big absolute swings, cancels in the mean
Gemini 3.7 All axes |Δ| ≤ 1.4 mostly cover 0 ≥0.25 Near-invariant

Axis heatmap

Honest takeaway: on 12–16 pairs per axis, few deltas clear a sign-test bar. The actionable signal is which models are volatile (Gemini 2.5 Flash, GPT mini, Grok) vs stable (Gemini 3.7, Haiku on flips), not a courtroom claim of systemic bias from p<0.05 alone.

Hiring decision flip rate (threshold = 70)

A twin “flips” when one résumé clears a shortlist cut of 70 and the other does not — same skills.

  • Gemini 3.7 / Haiku: 0% of twin groups
  • Sonnet 4.5 / Gemini 3.5 Flash: 4.1%
  • GPT-5.4 mini / Grok NR / Gemma 4 31B: 6.6%
  • Grok 4.20 Reasoning: 9.8%
  • Gemini 2.5 Flash: 14.8%

That is the recruiter-facing metric: not “mean MAE,” but would this person still make the shortlist if we only changed the header?

Flip rates

What this means for recruiters in LatAm

  1. Blind the header before LLM scoring (name, age line, city, education brand, maternity/disability asides) — or score twice and flag spreads.
  2. Pick models with low twin spread + low flip@threshold, not only low MAE. Gemini 3.7 led both accuracy and fairness here.
  3. Maternity and disability lines are not “noise” to the model — they move scores more than many surname cues. If your JD does not ask for them, strip them before screening.
  4. Prompt scale matters (v1 lesson): say integer 0–100 with examples, or you will measure scale chaos and call it bias.

Limitations

  • Synthetic CVs; gold = skill overlap, not human recruiter judgment.
  • 12–16 pairs/axis → wide CIs; few sign tests reject H₀.
  • Single run per model (no rerun consistency yet).
  • DeepSeek-R1, Qwen3 Next 80B and GPT-OSS 120B errored, so nine models are scored.
  • Spanish MX/CO/AR names only; not all LatAm ethnonyms.

What I’d measure next

  • Blind vs visible headers on the same v3 packs
  • Rerun consistency (3 seeds) for GPT mini / Grok
  • Spanglish JDs common in MX tech
  • Force fit_score through an integer enum tool

Where can we see it?


Alex Ramírez — Mexico City — AI for HR/recruitment — GitHub @davidramsem

Top comments (0)