DEV Community

Cover image for I built a hackathon judging app and it was too confident
Gaurav
Gaurav

Posted on

I built a hackathon judging app and it was too confident

DOGFOOD 2026 had a funny premise. Everyone builds the same thing: the portal that will judge the hackathon. The winner gets forked and used for real events. So you are building the thing that decides who gets the prize money, and then it gets used on you.

My entry is called Raptor Desk. Django, HTMX, SQLite, one docker compose up. This post isn't a feature tour. It's about the part that actually kept me busy: getting a fair ranking out of messy judge scores, then finding out my app was more sure of itself than it had any right to be.

Repo: https://github.com/GauravS13/raptor-desk

the official data is a trap

The organizers ship a fixtures file with the brief: 40 projects, 30 judges, 126 score sheets, each scoring functionality, quality and innovation. It looks boring. It isn't.

  • The plain average has a tie for first place. prj_11 and prj_34 both score 4.333.
  • prj_10 is 3rd on the average. It had two reviews, and one came from jdg_15, the second most generous judge in the set. Correct for that and it drops to 7th. That's prize money moving because of who happened to review it.
  • One judge (jdg_07) gives the exact same score on every criterion for every project.
  • prj_41 is prj_07 submitted twice.
  • Three different teams are called "StillTrail".
  • 49 reviews have no written comment at all.

The one that changed how I thought about the whole problem: judges differ more from each other than projects do. Judge leniency in this data runs from about −0.82 to +0.52. The gap between neighbouring projects is 0.02 to 0.10. So "who reviewed you" matters more than "how good you are" unless you correct for it.

the textbook fix breaks on this data

The first thing everyone reaches for is a per-judge z-score. Take each judge's scores, subtract their mean, divide by their standard deviation. Harsh and generous judges end up on the same scale.

On the official fixtures it gives NaN for 7 projects. Dividing by the standard deviation is the problem. The flat-lining judge has a spread of zero. So does any judge who only reviewed one project. Divide by zero, and those projects can't be ranked at all.

What I used instead is an additive model. Every score is the project's quality, plus the judge's leniency, plus noise:

score(judge, project) = quality(project) + leniency(judge) + noise
Enter fullscreen mode Exit fullscreen mode

Fit it with a little ridge regularisation so a judge with one review can't get a wild leniency estimate. Then alternate: estimate qualities given leniencies, then leniencies given qualities, until it settles.

q = (sum over judges of (s - b[judge]) + lam * mu) / (n_reviews + lam)
b = (sum over projects of (s - q[project])) / (n_reviews_by_judge + lam)
Enter fullscreen mode Exit fullscreen mode

It never divides by anyone's spread, so the flat-liner can't break it. That judge's scores carry no ranking information anyway, so the app detects it (identical scores on every criterion, every time) and leaves it out of the fit.

How do you know it's better? The fixtures have no right answer, so I made data that does. 200 fake hackathons where I know each project's real quality, with harsh and generous judges, noise, missing reviews and one flat-liner. Then I compare how well each method gets the true order back.

method how well it recovers the true order (Spearman) beats the plain average in
plain average 0.697 –
z-score with shrinkage 0.778 96% of runs
additive model 0.813 96% of runs

Good. Bias correction works. That part went fine.

then i made it say how sure it was

A ranking with no uncertainty is a lie with decimals. So the app resamples each project's reviews 300 times, refits, and counts how often each project lands in the top 1, 2, 3 and 5. That gives every project a "chance of a top-5 finish". Anything between 20% and 80% is flagged as a close call.

organizer results page with each project's chance of a top-5 finish

On the fixtures this says what you'd hope: places 1 and 2 look solid, places 3 to 5 are close calls. Then the app suggests which judge should review which close-call project to settle it, within a budget the organizer sets. Whatever is still close after that goes to a deliberation page where a person has to write down why they decided.

close calls with the judges the app suggests asking

I was pretty happy with this. Then I tested it.

then i checked if it was telling the truth

If the app says "90% interval", the true value should land inside it about 90% of the time. If it says a project has a 70% chance of a top-5 finish, projects like that should make the top 5 about 70% of the time. I had the 200 fake hackathons with known answers, so I could just check. 7,973 projects in total.

what the app said vs what happened: stated chance against real top-5 rate

It was too confident.

  • The "90% intervals" contained the true quality 52% of the time. Not 90.
  • For projects with one useful review it was 22%. With two, 48%. With three or more, 59%.
  • When the app said 95% or more, the project actually finished top 5 about 80% of the time.

The reason is kind of obvious once you see it. The resampling only shuffles the reviews a project already has. With two or three reviews there just aren't many ways to shuffle, so the spread looks small. It can't imagine a fourth judge who disagrees. With one review there's nothing to shuffle at all, so the interval has zero width. Very confident, about nothing.

It wasn't useless though. Two numbers made me keep it:

  • As a forecast it still clearly beats guessing. Brier score 0.073, against 0.109 for just saying the base rate every time (lower is better).
  • Of all the projects the ranking put on the wrong side of the prize line, half had already been flagged as close calls. The doubt was pointing at the right places, it just wasn't strong enough.

So I didn't hide it. The numbers are in the repo's JUDGING.md, generated by a script that CI reruns, so they can't quietly drift. And the app's job changed a bit in my head: chances are for pointing people at the doubt, never for deciding. A model never settles a prize place on its own.

The fix is known. You resample the judges' leniencies too, not just the reviews. I didn't ship it. It would have changed every chance the app shows, and I'd want to rerun the same calibration check on it first. Shipping a fix for overconfidence without measuring it would be the same mistake again.

one judge can pick the winner

The calibration result made me wonder what else the resampling couldn't see. It shuffles reviews but always keeps every judge. So what happens if you take a judge out?

I added a "kingmaker check" (the code is short). Refit the whole ranking once per judge, leaving that judge out. If removing one judge alone changes who's in the top 5, or who's first, flag it.

On the official fixtures:

without judge their reviews what changes
jdg_04 4 first place goes from prj_34 to prj_37
jdg_15 6 first place goes from prj_34 to prj_11
jdg_24 11 prj_16 and prj_18 move into the top 5, prj_33 and prj_37 drop out
six others 5 to 9 one project in or out of the top 5

9 of 30 judges each move a top-5 place on their own. And two of them flip first place, the one place the resampling called fairly safe. jdg_04 did four reviews. Four.

kingmaker check: judges whose removal alone changes the top 5

The reason is the same as before. Each judge's leniency is estimated from a handful of reviews. Take a judge out and the corrections for everyone they overlapped with shift a little, and at the top the gaps are tiny.

I thought about automatically down-weighting these judges and decided not to. A kingmaker isn't a bad judge. It just means a prize is resting on one person's opinion. So the app shows the list next to the close calls, and the organizer can ask for more reviews or decide on the record. There's also an API for it: GET /api/events/{id}/kingmakers.

three bugs i actually learned something from

1. signed by itself. Results get frozen into a snapshot signed with Ed25519, and anyone can verify it. My first verifier checked the signature using the public key stored inside the document. Spot it? Anyone can change the results, sign them with their own key, put their key in the document, and it verifies perfectly. It's self-consistent, it's just not mine. Fixed on Sunday night:

# before: trusts whatever key the document brings
return signing.verify(payload, signature["signature"], signature["public_key"])

# after: the key itself has to be this deployment's
if signature.get("public_key") != public_key_hex():
    return "signed with a different key, not by this deployment"
Enter fullscreen mode Exit fullscreen mode

The offline verifier now takes the portal's key from outside the document too.

2. the audit log lost exactly the attacks. Every vote attempt is audited, accepted or refused. Refusals (voting twice, voting for your own team) raised an error inside the database transaction. The error rolled back the transaction, and the audit entry with it. So the log had every honest vote and none of the cheating attempts. Simplified:

# before
with transaction.atomic():
    if refused:
        audit.record("vote.refused", ...)   # rolled back with everything else
        raise refused_error

# after
with transaction.atomic():
    refusal = check(token, project)
    if refusal is None:
        ...create the vote and audit it...
if refusal is not None:
    raise refused(request, token, *refusal)  # audited outside the rolled-back block
Enter fullscreen mode Exit fullscreen mode

3. same data, different numbers. The resampling uses a fixed seed, so I assumed it was reproducible. It wasn't. The reviews came out of the database in whatever order the database felt like, and with a seeded random generator a different input order means different resamples. Same scores, different results: prj_37's chance of a top-5 finish was 0.64 with the reviews in one order and 0.58 in another. The docs computed from the fixtures file and the app read the same scores from the database in a different order, so they disagreed. Fix was one line, sort the input before doing anything. Then CI regenerates the docs from the engine and fails if they don't match.

the spec was harder than it looked

The organizers' acceptance checker only tests tiers 1 and 2: public gallery, judge isolation, CSV export. Tiers 3 and 4 (public voting, webhooks, an embed, import and export) aren't checked by anything.

So I wrote my own checker for those, same report format, and said in the README that it's mine. The thing I learned writing it: a portal that refuses everything passes every "should be refused" test. If you only check that a participant gets 403, a broken app that 403s everyone looks perfect. So every "refused" check has an "allowed" twin right next to it. A vote for your own team is refused, and someone from another team voting for the same project goes through. If both don't hold, the check fails. 32 checks in total.

Smaller ones that bit me:

  • My own security headers broke the API docs. A strict Content-Security-Policy blocked the Swagger UI from a CDN, so /api/docs was a blank page. Serve it locally.
  • One route ate another. /api/events/import got matched as /api/events/{event_id} with an event id of "import". Moved it to /api/event-bundles.
  • Imports went straight to public. An exported event that was published came back in as published, so importing an archive made it public instantly. Imports now land as drafts.
  • My CI was never green. I'd pinned astral-sh/setup-uv@v10. That version doesn't exist. Meanwhile the README was going on about proof. Now CI also reruns the official checker against a fresh docker compose up and compares its output with the committed report byte for byte.

what i cut, and what i'd redo

Cut, and fine with it (every one of these, with the reasons, is in the decision log):

  • PostgreSQL as the default. SQLite in WAL mode with immediate transactions is plenty for a hackathon and keeps it one command.
  • Celery and Redis. A small outbox table and a worker container do the emails and webhooks.
  • Alpine.js. Its normal build needs unsafe-eval, which my CSP refuses.
  • Pairwise "A or B?" judging replacing the scores. It's there, but beside the scores, not instead of them. A rubric with weights and a pass/fail gate can't be expressed as pairs.
  • Blocking votes per IP address. A room full of people on the venue Wi-Fi looks exactly like ballot stuffing. Organizers can mark trusted networks instead, and suspicious bursts get held for review, never thrown away.
  • Sealed commit-reveal voting. Didn't have time to do it properly.

Would redo:

  • The score ledger from the first commit. I added a hash-chained, signed ledger of every score after reviews already existed, so I needed a backfill migration for the old ones. It works, but "verifiable from the start" beats "verifiable after a migration".
  • The resampling. See above. Resample judges, not just reviews, and check calibration again before shipping it.

try it

git clone https://github.com/GauravS13/raptor-desk
cd raptor-desk
docker compose up
# http://localhost:8080 , demo accounts are on the sign-in page

python3 tools/run.py .dogfood.toml --fixtures data/fixtures.json   # official checker, 7/7
uv run python tools/simulate_proof.py --check                     # every number in this post
Enter fullscreen mode Exit fullscreen mode

The calibration and kingmaker numbers come from tools/simulate_proof.py, and the method is in JUDGING.md. If you think my maths is wrong somewhere, I'd honestly like to know.

Built for #DogfoodHackathon by Hackathon Raptors.

Top comments (0)