This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
I Built an AI Judge That Doesn't Decide Who Wins
My friend participates in hackathons frequently and builds strong projects, but she had an important concern:
How can you know that judges are applying the same rubric consistently?
A project can receive different scores from different judges even when everyone is evaluating the same criteria.
So I built FairJudge AI.
FairJudge is an AI-assisted hackathon evaluation system designed to make judging more transparent, evidence-based, and consistent.
It can:
- Evaluate projects against a supplied rubric
- Connect scores to supporting evidence
- Identify missing evidence
- Explain why a score was given
- Compare multiple judges
- Detect significant scoring disagreements
- Recommend human review when judges disagree
But there is one important thing FairJudge doesn't do:
It does not decide who the objectively correct winner is.
The final decision always remains with human judges.
The Problem
Hackathon judging involves multiple judges evaluating projects across several criteria.
Disagreements can happen because:
- Judges interpret criteria differently
- Evidence may be incomplete
- Some claims cannot be verified
- Judges may emphasize different aspects of the same project
Instead of building another AI that simply says:
"Project X should win."
I wanted to build something more transparent:
Why was this score given? What evidence supports it? What evidence is missing? Where do judges disagree?
That became FairJudge AI.
Demo
🎥 Demo Video:
https://claude.ai/artifact/9JLZQKa7C33ADuJg2Qk5wC
The demo shows:
- Evaluating a hackathon project
- Analyzing supporting and missing evidence
- Generating an AI-assisted evaluation
- Comparing multiple judges
- Detecting significant disagreements
- Rejecting unsupported evidence
- Recommending human review
How FairJudge Works
The core workflow is:
text
Project + Rubric
↓
Evidence-based AI Evaluation
↓
Evidence Validation
↓
Deterministic Python Scoring
↓
Multiple Judge Comparison
↓
Disagreement Detection
↓
Human Review
Top comments (0)