DEV Community

Cover image for The judges agreed no better than chance: 4 lessons from building a hackathon judging platform
Tushar Agarwal
Tushar Agarwal

Posted on

The judges agreed no better than chance: 4 lessons from building a hackathon judging platform

A five minute demo of Quorum, the self hosted hackathon platform this post is about. The code is open: github.com/TusharTechs/quorum.

For DOGFOOD 2026 the brief was simple to say: build a self hosted hackathon platform. Registration, teams, submissions, judging, voting, results. Everyone got the same fixture to test with: 41 submissions (one of them a duplicate), 30 judges and 123 reviews across 8 tracks.

Before building a single screen, I measured how much those judges agreed with each other.

ICC(1) came out at minus 0.006. A permutation test with 2,000 shuffles put p at 0.504. In plain words: these judges agreed with each other no better than chance. Any podium printed from this data is noise with a trophy on it.

That one number changed what I was building. Every platform can print a ranking. The useful thing is to say which parts of the ranking the data actually supports, and to settle the rest fairly and on the record. It became the tagline: every project judged, every tie decided, every team answered.

Here are the four things that taught me the most.

1. The architecture decision that shaped everything: fail loudly in the app, enforce invariants in the database

The security core of a judging platform is judge isolation: a judge must never see another judge's scores. The obvious tool is Postgres row level security. I wrote it into the spec, and then deliberately did not build it.

Row level security filters rows silently. That is exactly what you want for a list page, and exactly what you do not want for maths. A preview ranking, or the code that picks the next pair of projects to show a judge, aggregates over other judges' data. Run that under a judge's session with row level security on and it does not fail. It quietly computes a wrong answer from a subset of the rows. For a judging platform, a wrong number that looks right is the worst failure there is.

So the rules split in two:

  • In the application, and loud. Every route has to declare a policy with a decorator, and the app refuses to boot if one does not. Scores can only be read through a handful of scoped repository functions, and a lint test fails the build if any other module queries the score tables directly. A judge asking for someone else's scores gets a plain 403.
  • In the database, for everyone. Whatever must hold no matter which code path runs lives in Postgres triggers: the audit log is append only and hash chained, ranking runs are immutable, scores must sit inside the criterion's scale, and submissions freeze at the deadline:
IF now() < closes THEN RETURN NEW; END IF;
IF TG_OP = 'INSERT' THEN
  RAISE EXCEPTION 'deadline_passed: submissions closed at %', closes;
END IF;
Enter fullscreen mode Exit fullscreen mode

The same thinking decided the infrastructure. Postgres is the only moving part besides the app. Emails and webhooks go into an outbox written in the same transaction as the change that caused them, background jobs use SKIP LOCKED, rate limits are ON CONFLICT counters, and web replicas boot one at a time behind an advisory lock. No Redis, no Celery, four containers, and one command: docker compose up.

Quorum architecture

How do I know isolation holds? An authorization matrix calls all 96 API operations as 6 different roles. A canary crawl plants secrets (a private draft title, a judge's private note) and checks they never appear in any response, CSV or export. And a live probe fires 344 cross role requests at the running stack. Zero leaks.

2. The scoring algorithm that looks right and is broken

When judges differ in how generous they are, you need normalization. The textbook fix, and the one most platforms ship, is per judge z scores: subtract each judge's mean and divide by their standard deviation. On this fixture it breaks in three ways:

  • It divides by zero. One judge gave every single project a 4. Their standard deviation is zero. Two more judges have only one review. z scores are undefined for all three.
  • It erases information. A judge with two reviews always produces +0.71 and minus 0.71, whether they scored 4.9 and 5.0 or 1 and 5.
  • It punishes bad luck. A judge who happened to get a strong batch looks generous, so every project in that batch gets pulled down.

I tested it rather than trusting intuition. In a simulation built on the fixture's exact assignment design, with a known true quality for every project, z scores did worse than the plain raw average at a moderate spread of judge leniency, and picked the true winner 23% of the time against 29% for the raw average.

What Quorum ships instead is a random effects model: each judge gets a leniency offset, shrunk toward zero by an amount that REML estimates from the data itself. When judges really differ, it beats the raw average (Kendall τ up 0.018 at a leniency spread of 0.4, up 0.070 at 0.8). When they do not, the shrinkage grows and it collapses back to the raw average, so it costs nothing. The judge who gave everything a 4 gets weight zero by a rule published before anyone registered, and every project page shows exactly what that did.

Why Small Relay is #30

Small Relay was 12th on the raw average. One of its two judges was the all 4s judge. Excluded by rule, it lands 30th, and the page shows the arithmetic line by line.

Then came the insight I did not expect: no normalization can fix a bad assignment. If a group of projects is always reviewed by the same group of judges, each judge's offset shifts all of those projects equally, so the calibrated score is provably identical to the raw average. In simulation, judges in separate panels gave exactly zero gain; overlapping batches gave +0.036. A harsh panel and a weak batch of projects look the same unless batches overlap. Your assignment algorithm is part of your scoring algorithm, so Quorum's planner maximizes overlap and shows how connected the judges are before you commit a single batch.

And the bug I nearly shipped: my first focus round planner, which sends spare judge time to the projects whose prize is still in doubt, only looked at the overall podium. The fixture also pays a Best in track prize in each of its 8 tracks. A project that was hopeless overall but a coin flip for its track prize got no attention at all. The fix scores a project's uncertainty on both prizes, overall and in its track, and only spends reviews where either is still open.

How certain is the podium

Finally, when the data cannot separate projects at a prize line, Quorum does not guess. It opens a head to head round, defined before the event, with three judges who have no conflict with the tied projects. Their comparisons decide the order, and if even those are not decisive, the organizer records a decision with a written reason in the audit log.

3. The Docker gotchas

os.cpu_count() lies inside containers. My default was one gunicorn worker per core, capped at 8. Sensible on a laptop. When I deployed the public demo to a host with a 1 GB memory limit, the container reported the host machine's cores, so it would have started 8 workers. I measured each worker at about 85 MB idle and 170 MB once the local AI model loads: roughly 1.4 GB in a 1 GB box, killed on the first search. The fix reads the limits the container actually has:

def default_workers():
    n = min(os.cpu_count() or 2, 8)          # the host's cores, not yours
    quota, period = open("/sys/fs/cgroup/cpu.max").read().split()
    if quota != "max":
        n = min(n, max(1, round(int(quota) / int(period))))
    mem = open("/sys/fs/cgroup/memory.max").read().strip()
    if mem != "max":
        n = min(n, max(1, int(mem) // (250 << 20)))   # about 250 MB per worker
    return max(1, n)
Enter fullscreen mode Exit fullscreen mode

Behind a proxy, every request comes from a private address. Two bugs, one root cause. The internal metrics page was restricted to private networks, which was correct until Caddy (or a hosting platform's edge) sat in front: then every request, including the whole internet's, arrived from a private address, and the metrics page was public. And rate limits read the client's address from the leftmost entry of the forwarded header, which is whatever the client chose to send, so a forged header dodged per network limits. The fixes: refuse the metrics page whenever a forwarding header is present, and trust only the rightmost hop that is not your own proxy. A load test with 300 voters from distinct networks is what exposed the second one.

Three replicas, three migrations. Scaling to three web replicas meant three containers racing to migrate the database at boot. A Postgres advisory lock around the boot step fixed it: the first replica does the work, the others find nothing left to do.

Proving "offline" honestly. The rules say the platform must work with the network off. Instead of trusting that, the check runs inside a Docker network marked internal, first proves it cannot reach the internet, and only then runs the official checker (7 of 7), our tier 3 and 4 checker (21 of 21) and the isolation probe. The local AI model ships inside the image, so it runs in there too.

4. The spec line I thought was simple: "a closed event refuses submissions"

One line in the spec, one request in the official checker. It took more thought than the calibration model.

  • Refused for the right reason. A late submission that is also missing a field must fail with deadline_passed, not a validation error, or you have just told a latecomer to go fix their form. So the order is fixed: authenticate, check the role, check the deadline against the server's clock, and only then validate the input.
  • Frozen, not just blocked. After the deadline a team can still read its submission, but the content is frozen. The trigger above compares every content column, so not even a database shell can edit a project after close, and the submission is sealed into the audit chain with its content hash.
  • Refusals must not leak either. The same goes for "a judge cannot see peer scores". While recording the demo video, I noticed that asking for a real judge's scores and asking for a judge who does not exist returned 403s with different wording. Same status code, but the difference told an outsider which judge IDs exist. Now the two refusals are identical, and the test compares the full response bodies, not just the status codes.

Lesson: test what a refusal says, not only its status code.

5. AI in a judging platform, without letting it make things up

A judging platform that invents facts is worse than one with no AI, and the rules ban external APIs anyway. So Quorum runs a 23 MB sentence embedding model on the CPU inside the stack and uses it only to choose, never to say:

  • "Who hasn't started?" is matched by meaning to one of 15 named skills. The skill runs the same permission checked query as the page, and the answer says which skill ran and what it read.
  • A judge asking "who is winning?" gets nothing, because judges are not allowed to know that.
  • A feedback coach checks whether a judge's comment has a concrete next step and which criteria it covers. It advises. It never scores.

Ask Quorum

What I would tell the next person building one

  • Measure how much your judges agree before you design the results page. If they agree at chance level, the product is the uncertainty.
  • Prefer loud failures in code and invariants in the database over silent filters.
  • Your assignment algorithm is part of your scoring algorithm.
  • Read the container's limits, not the host's.
  • Test what a refusal says, not only its status code.

The numbers, for the skeptical: official checker 7/7, our tier 3 and 4 checker 21/21, 188 tests on real Postgres, 344 isolation probes with 0 leaks, a load test with 300 voters on 3 replicas and 0 lost votes, and every published ranking recomputes byte for byte with the operating system's own Python: MATCH.

GitHub logo TusharTechs / quorum

Self hosted hackathon platform that judges fairly: calibrated scores, decided ties, feedback for every team, and a local AI assistant that never makes things up. Runs offline with one command.

Quorum

Every project judged. Every tie decided. Every team answered.

Demo video · Live demo · Run it · Five minute tour · Architecture · Judging method · Normalization proof · Threat model · Verification · Operations · Demo script

Watch the five minute Quorum demo

For judges: you want to… Go to
Watch it in five minutes Demo video: one full lifecycle, create, submit, judge and publish, with each scene mapped to a judging criterion
Try it now, nothing to install Live demo: one click to be the organizer, a judge or a participant. It resets every hour and sends no e-mail
Run it yourself docker compose up: seeded with the official fixture, works with the network off
See every tier working What is built · official checker, 7/7 · T3/T4 checker, 21/21
Check the judging maths JUDGING.md · normalization proof: raw vs normalized scores, rank changes, the method defended
Check security
…

Top comments (0)