For DOGFOOD 2026, Hackathon Raptors gave teams an odd brief: build the platform they'll run their own hackathons on. Registration, teams, submissions, judge assignment, scoring, results, certificates, an archive of past events. We're Team CodeHawk, and our entry is HackFlow, a self-hosted platform that comes up with a single docker compose up.
I could walk you through the features, but the parts worth writing about are the four places where something we thought was simple turned out to be wrong. Three of them sit in the code that decides who wins.
Repo: https://github.com/sarvan-2187/dog-food
One database, on purpose
The decision that shaped the rest of the project was keeping everything in one place: one FastAPI process and one Postgres database. The Docker build has two stages. Node builds the React app and then gets thrown away, and the slim Python image that's left serves both the API and the frontend. At runtime there are two containers in total.
It would be easy to read that as a shortcut, so here's the reasoning. A hackathon platform has to hold rules like these:
- nobody saves a submission after the deadline
- a rubric can't change once a judge has scored against it
- one team per person per event
- one vote per person per project
Each of those is a check that needs to see consistent data at the exact moment someone writes. If teams, submissions and scoring lived in separate services, every rule would turn into a distributed consistency problem. With one database, each one is a query inside a transaction, and usually a unique index backs it up. Almost every later decision got easier because we made this one early.
The scoring algorithm we almost shipped
Judges don't grade the same way. One gives everything a 90, another treats 6 out of 10 as high praise. If you average raw scores, the projects that happened to draw the generous judge win, and that's luck.
So HackFlow converts each judge's totals into z-scores against that judge's own mean and spread, then averages them per project. A score stops meaning "8 out of 10" and starts meaning "well above what this judge usually gives."
That's the textbook version, and we had it working. Then, on the day of the deadline, we audited our own implementation against the official fixture data (40 projects, 30 judges, 123 scores) and found it was wrong in ways none of our tests had caught.
To be clear about the stakes: this was all test data. HackFlow hadn't run a real event yet, so no published result was ever wrong, and every fix below is in the version we submitted.
Bug 1: the weights didn't mean what the organizer typed
We calculated the raw total as Σ weight × value. That works as long as every criterion uses the same scale. Now picture an organizer who sets Technical at 0 to 10 with a weight of 70%, and Presentation at 0 to 100 with a weight of 30%. Presentation ends up carrying about 81% of the total.
Nothing throws an error, and the screen still says 70/30, but the ranking is following a split nobody chose. The fix was to scale each value by its own maximum before weighting, then divide by the sum of the weights. When every criterion shares a scale, the new formula gives exactly the same answer as the old one. That's the uncomfortable part: all our fixtures used one scale, which is why our tests never noticed.
Bug 2: judges with nothing to say were dragging scores down
A judge who scored only one project, or gave every project 4/4/4, has a spread of zero. We'd avoided the divide-by-zero by giving those judges a z-score of 0, and then we averaged that 0 in with everyone else's.
That pulls a project's score toward zero in proportion to how many of those judges it happened to draw. Two projects with the same real evidence could rank differently purely because of who was assigned to them. On the fixture data, 12 projects were in the wrong position because of this. Now a judge whose scores say nothing about relative order is left out of the average entirely, instead of being counted as a vote for "average."
Bug 3: floating point invented an opinion
We checked for zero spread with σ == 0. Weighted totals produce values like 0.1 × 3 and 0.3, which aren't equal in floating point. So a judge who gave identical marks looked like they had a tiny spread, and dividing by that tiny spread turned rounding error into z-scores of +1 and -1. That's a full point of normalized score, out of nothing. We now compare against a relative tolerance, σ > 1e-9 × max(1, |μ|).
There was a smaller one hiding in our seed data, too. The rubric weights were 0.3333, 0.3333 and 0.3334, which was enough to split real ties by a ten-thousandth. Two projects tied at 4.33 showed up as first and second. Exact thirds fixed it.
What changed once it was right
With the bugs fixed, normalization moved the standings more than we expected. Only 4 of the 40 fixture projects kept the rank a plain average would give them, and 9 of the top 10 moved. That movement is the whole point of normalizing. It shows how much judge harshness was shaping the raw order. Dry Harbour went from 31st on raw scores to 9th. Its judges were tough on everyone, and it was one of the best things they saw.
The lesson I'd pass on: a test that recomputes your formula with your formula can't tell you the formula is wrong. The test that finally earned its keep rebuilds the standings from the raw fixture file using a separate implementation, then checks the live results endpoint against it. Every bug above now has its own test, written to fail on the old behavior, so none of them can come back quietly. We trust the rankings more now than we did before the audit, because we've watched them fail and know why.
The Postgres lock that hung our test suite
We didn't use a migration framework. Instead, a few idempotent steps run every time the app boots, so running docker compose up on an old database upgrades it in place. The obvious way to add a column looks harmless:
ALTER TABLE users ADD COLUMN IF NOT EXISTS session_version integer DEFAULT 0;
It looks safe to run on every boot, and it isn't. Postgres takes an exclusive lock on the table before it checks whether the column already exists. Our test suite keeps transactions open, so the app's startup sat waiting for that lock and the whole suite hung. There was no error message, just nothing happening.
The fix is to ask information_schema whether the column exists first, and only run the ALTER when it's actually missing. We applied the same caution to a new unique index enforcing one team per person per event: it only gets created if the existing data already satisfies it. If it doesn't, the app still boots, logs the duplicates, and lists them on the admin page so a person can sort them out. We didn't want a boot step that could fail on old data to take an event offline.
Disqualification is harder than hiding a row
The brief says organizers can disqualify an entry. We assumed that meant hiding it from the results.
Normalization makes that wrong. Say a judge gave the disqualified entry a 1 out of 10. That score is part of the judge's mean. If you only hide the row, the 1 keeps dragging the judge's mean down, and every other project that judge scored looks a little better than it should. Removing one entry would quietly change the scores of projects that had nothing to do with it.
So HackFlow drops the entry's scores before normalization runs, and recomputes every judge's mean and spread as if that entry had never been judged. The scores themselves are kept, so if an organizer reinstates the entry, the next read of the results puts everything back.
The same reasoning decided how the judging deadline works. It's soft on purpose: a late score is recorded and flagged, never refused. Normalization can correct for a harsh judge, but it can't make up a review that never happened.
What we haven't solved
Two limits are still there, and we'd rather name them than have someone find them.
A judge who scored exactly two projects always produces +1 and -1, whether the gap was 4.9 against 5.0 or 1 against 5. That comes with per-judge z-scores on small samples. The real remedy is giving each judge more entries, which is what the assignment step's judges-per-submission setting is for.
Normalization also only removes a judge's general harshness. It does nothing about a judge who favors one particular project. Catching collusion needs a different tool, and for now that's the per-judge CSV export and the audit log.
Stack and credits
Python 3.12, FastAPI, SQLModel, PostgreSQL 16, React 18 with TypeScript, pytest and Playwright, all run with Docker Compose. We built HackFlow with help from Claude Code, Anthropic's AI coding agent, and say so in the README. The decisions are ours, and so are the bugs.
What we ended up with is a platform that runs a whole event, from sign-up to certificates and an archive, from one docker compose up on seeded data, and it passes all 7 of the official DOGFOOD checks. The judging part is the piece we're proudest of, mostly because of how much it taught us.
If you're building anything that turns people's scores into a ranking, test it with criteria on different scales and with a judge who gives everyone the same mark. Those two cases would have caught our two worst bugs.
Team CodeHawk



Top comments (0)