Building a judging platform in 72 hours: the maths that fought us, a role-isolation bug, and the feature we cut
For Dogfood 2026, Hackathon Raptors gave every team the same assignment: build the open-source, self-hosted submission and judging platform they would actually run their future events on. Our entry is RaptorGate - Django 5.2, PostgreSQL 16, one docker compose up to a seeded portal on localhost. We claimed tiers T1 and T2, verified 7/7 by the official checker, with our T3/T4 work disclosed as partial slices rather than claimed. This is what building it was actually like.
The spec was harder than it looked
We lost our first build to a calendar. The organizers moved the event window by 24 hours, and rule 04 says all project code must be written inside the window. Our first repo - over a hundred commits, checker already passing - predated kickoff. Copying that code into a fresh repo would have walked straight into disqualification, so we threw it away and rewrote the whole thing in-window. Same architecture, new code. That decision hurt more than any bug.
The second moment was deadline enforcement. The checker's closed-submission probe expects a 403 when it POSTs after the deadline. But Django's CSRF middleware can produce that same 403 before your deadline logic ever runs. A probe that passes for the wrong reason is not a pass, so we verified the deadline handler separately with CSRF disabled. That one is now burned into how I read any test result.
The schema I would redo
Our models enforce the important boundaries in endpoint queries and form validation, not in the database. A project's team and track are not constrained to the same event at the DB level, and an assignment's judge and project are not either. Every exposed path checks, and the tests cover the denial cases - but one future write route that forgets a check writes cross-event garbage into storage. Score criteria are JSON without a matching model-level constraint, for the same reason.
If I redid the schema: event-scoped foreign-key constraints in the database itself, and a real constraint on criteria shape. Application-level checks are where bugs go to wait.
The normalization maths that fought us
Judges grade on different scales, so raw means across judges are not comparable. Our method: weight each review (0.40 functionality, 0.35 quality, 0.25 innovation), then per judge compute the mean and population standard deviation over their valid reviews, and adjust each score to clamp(3 + (r - mu) / sigma, 1, 5). A judge with zero variance - our fixture has one with three identical reviews - contributes a neutral 3.0 instead of invented differences.
Then we built the proof, and the proof fought back. On the official fixture, 32 of 40 projects change position after normalization. One climbs 22 places, another drops 16. Our first reaction was "it works." Our second was "wait - is the new ranking more accurate, or just different?" The honest answer is that we cannot claim more accurate. Assignments are disjoint, some judges reviewed a single project, and centering on one review is just a neutral 3. The maths is defensible; what it proves is limited. So we shipped it with that caveat printed next to the standings, and told organizers to check raw scores and review counts before treating rank movement as a prize decision. The proof is reproducible from the repo: one management command, byte-for-byte stable output, same code path as the live standings.
The role-isolation bug
Judge invitations go out by email. An early version of our invitation path keyed invited users globally by email. Here is the attack: someone is already a judge in event A. An organizer of event B invites the same email. If the system reuses the global invited account, whoever holds event B's accept token can replace that account's password - and seize access to event A.
Role isolation failed not inside one event but across two. We fixed it fail-closed: the invitation path now rejects any email already tied to a judge in any event, with a regression test named after the hijack. A proper long-term fix needs proof of account control without a password reset. We found this in review, not in production, which is exactly where you want to find it.
The same paranoia shaped the rest of the judging paths: the score-write route re-checks track membership, assignment and own-team conflict even though assignment creation already checked them, and a judge requesting a peer's scores through the API gets a 403 the checker probes for.
The feature we cut: pairwise judging
Pairwise comparison is the fashionable answer to judging noise, and we cut it without regret. With the hours left, a half-built pairwise mode would have been a broken claim on the tier ladder, and this event scores you on what you claim. Weighted scores plus documented normalization, actually verified, beat a pairwise mode nobody had time to test. The repo says plainly: no pairwise mode. Choosing what not to claim turned out to be worth real points.
What the last hour cost us
We shipped the interactive public demo minutes before the freeze with no test coverage on it, and it showed: two bugs (a comment endpoint returning 405, a vote-submission JSON error) were found by clicking, not by our suite. The submission form itself went in about two minutes past the stated freeze. It was accepted, but that was luck, not execution. Tiers first, polish last - we learned it the expensive way.
Numbers
- 296 Django tests and 299 pytest tests on the submitted source, plus the official checker's 7/7 on the claimed tiers
- Official fixture seeded by one command: 8 tracks, 30 judges, 41 project records, 126 score rows
- Normalization proof:
python src/manage.py normalization_proof --event evt_01 - Repo: https://github.com/muffedd/raptorgate-dogfood-2026
Thanks to Hackathon Raptors for a spec with teeth. If you ran the same weekend, I would genuinely like to read how your normalization went.
Top comments (0)