DEV Community

Cover image for I Built an AI Judge That Doesn't Decide Who Wins.
DIKSHIT GARG
DIKSHIT GARG

Posted on

I Built an AI Judge That Doesn't Decide Who Wins.

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

I Built an AI Judge That Doesn't Decide Who Wins

My friend participates in hackathons frequently and builds strong projects, but she had an important concern:

How can you know that judges are applying the same rubric consistently?

A project can receive different scores from different judges even when everyone is evaluating the same criteria.

So I built FairJudge AI.

FairJudge is an AI-assisted hackathon evaluation system designed to make judging more transparent, evidence-based, and consistent.

It can:

  • Evaluate projects against a supplied rubric
  • Connect scores to supporting evidence
  • Identify missing evidence
  • Explain why a score was given
  • Compare multiple judges
  • Detect significant scoring disagreements
  • Recommend human review when judges disagree

But there is one important thing FairJudge doesn't do:

It does not decide who the objectively correct winner is.

The final decision always remains with human judges.

The Problem

Hackathon judging involves multiple judges evaluating projects across several criteria.

Disagreements can happen because:

  • Judges interpret criteria differently
  • Evidence may be incomplete
  • Some claims cannot be verified
  • Judges may emphasize different aspects of the same project

Instead of building another AI that simply says:

"Project X should win."

I wanted to build something more transparent:

Why was this score given? What evidence supports it? What evidence is missing? Where do judges disagree?

That became FairJudge AI.

Demo

🎥 Demo Video:

https://claude.ai/artifact/9JLZQKa7C33ADuJg2Qk5wC

The demo shows:

  1. Evaluating a hackathon project
  2. Analyzing supporting and missing evidence
  3. Generating an AI-assisted evaluation
  4. Comparing multiple judges
  5. Detecting significant disagreements
  6. Rejecting unsupported evidence
  7. Recommending human review

How FairJudge Works

The core workflow is:


text
Project + Rubric
       ↓
Evidence-based AI Evaluation
       ↓
Evidence Validation
       ↓
Deterministic Python Scoring
       ↓
Multiple Judge Comparison
       ↓
Disagreement Detection
       ↓
Human Review
Enter fullscreen mode Exit fullscreen mode

Top comments (0)