Built for the @partnerships_raptors Hackathon Raptors DOGFOOD hackathon: 72 hours, build the submission and judging platform that will judge you. FastAPI, PostgreSQL 16, Next.js, one docker compose up. Code: github.com/broforce6909-cmd/DogFood.
My judging docs contained a table explaining why you must not put an epsilon in the denominator when you normalize a judge who gave every project the same score. It said the judge would "become the most influential voice in the event."
After the build, I tested that sentence on the organizers' own data. It is false. The thing that actually mattered was something I had filed under "small detail." This post is about that, and about three other places where the problem was harder than the spec made it sound.
1. Normalizing judges: what actually moves the ranking
The setup. Judges do not share a scale. In the organizers' fixture, judge averages run from 2.0 to 4.22. I standardise each judge (z = (score - judge_mean) / judge_sd), average z-scores per project, and rescale to 1-5. Two edge cases needed rules: a judge with zero spread, and a judge with very few ballots.
My rules: a zero-spread judge maps to z = 0 ("I am not distinguishing between these"), and every judge's mean and sd are shrunk toward the field's with a constant k = 3.
The test. I ran the real scoring.py on the organizers' fixtures.json (30 judges, 126 ballots, 41 projects before I dedupe a duplicate team; equal rubric weights assumed) and compared alternatives. Ranking agreement with what I shipped, as Kendall tau:
| Method | tau | Top-5 overlap |
|---|---|---|
| Raw mean, no normalization | 0.80 | 4/5 |
| Plain z-score, epsilon in the denominator | 0.80 | 4/5 |
| Shrinkage on, but a flat judge borrows the global sd | 0.92 | 5/5 |
| Shipped: flat judge = z 0, shrinkage k = 3 | 1.00 | 5/5 |
Finding 1: normalization matters. 13 of 41 projects move three or more places versus the raw ranking, and raw third place is not normalized third place. Averaging raw scores across judges who use different scales gives a measurably different leaderboard.
Finding 2: my epsilon argument was wrong. An exactly flat judge has score - mean = 0, so 0 / (0 + 1e-9) = 0. The epsilon version produced the same ranking as no shrinkage at all. It did not blow up. And the "flat judge" rule versus borrowing the global sd moved one project in the top ten. The rule is still the principled choice, but I had sold it as a disaster-prevention measure and it is not one on this data. Only 3 of 30 judges are flat, and two of those have a single ballot, which is flat by definition.
Finding 3: the real hazard was the small-n judge. For a judge with two ballots, population sd gives z = +1 for the higher and z = -1 for the lower, always, whether the two scores differ by 0.33 or by 1.0. In the fixture, 8 of 30 judges have two or fewer ballots, and the six two-ballot judges' gaps range from 0.33 to 1.0. Without shrinkage each of them hands out the same-sized push regardless of how much they actually discriminated. That, not the flat judge, is what the plain z-score gets wrong (tau 0.80 again), and shrinkage is what fixes it.
Finding 4, the disappointing one: k is a hand-picked constant and the top five is sensitive to it.
| k | Kendall tau vs k = 3 | Top 5 |
|---|---|---|
| ~0 (no shrinkage) | 0.80 | one project in, one out |
| 1-5 | 0.94-1.00 | identical |
| 10 | 0.94 | third place drops out of the top five |
| 20 | 0.92 | further reshuffle |
| ∞ (everyone is average) | 0.89 | ... |
There is a stable plateau around k = 1 to 5, and I picked k = 3 inside it. But there is no ground truth here, so I cannot tell you k = 3 is right, only that it is not fragile nearby. It is published in the API response as shrinkage_k. The honest claim is "inspectable and stable in a range," not "optimal." I chose a fixed constant over a hierarchical model because with a dozen judges and a few ballots each, the between-judge variance would itself be estimated from almost nothing. You would get a more impressive number with no more information in it.
Reproduce all of it with normalization_sensitivity.py in the repo (no database needed).
A related catch. My first version z-scored against one global pool. A judge restricted to one track was being compared against a mean that included tracks they never saw, so a track that simply drew stronger projects looked like a harsh judge, and normalization cannot tell those apart. normalize_by_track() now partitions by the submission's track (not the judge's, since an unrestricted judge spans tracks). A normalized 4.1 in one track and a 4.1 in another are still not the same claim, and I wrote that into the docs rather than hide it.
2. Role isolation: make the bad request unexpressible
The rule is one sentence: organizers may read every ballot and may not write one. Everywhere else, organizer and admin can do anything to an event. Scoring is the single exception, because anyone who can author a judge's score can forge a result, and nothing downstream (export, normalization, audit log) could tell a forgery from a judgement.
So PUT /api/judging/assignments/{id}/scores has no "on behalf of" parameter at all. Refusing a request is a check. Not being able to express it is a property. Checks drift; properties do not.
The unit of isolation is the ballot (a judge_assignments row), not the user or the submission. Three requirements collapse into one ownership check: a judge cannot see a peer's scores, a track judge cannot see another track, and scoring closes outside the judging window. A Score has no rules of its own, so it is readable exactly when its ballot is.
One small decision worth copying: a peer's ballot returns 403, not 404. That a peer holds a ballot is not a secret (the organizer grid shows it), so pretending otherwise protects nothing. Another team's draft returns 404, because there the existence is the secret.
3. Check every feature's read side against its write side
After all four tiers worked, I audited with one question per feature: can a real user produce the state this code reads? Every one of these passed its own tests:
-
Voting configuration had no write path.
voting_method,vote_credits,votes_per_voterandvoting_accesswere read by every voter route. NeitherEventCreatenorEventUpdatehad a field for them. A real organizer could not turn on quadratic voting; every event that used it was seeded. -
Six of ten webhook topics never fired.
GET /webhooks/topicsadvertised all ten. Six were never passed tohooks.schedule()anywhere. - Certificate lookup had no auth check. Any participant could list any project's codes.
- CSV import skipped the stored-XSS validation that the JSON API had. Two write paths to one column, one validator.
The pattern is that I had tested what I built, not what a user could reach. Demo data hid all four.
4. The benchmark that lied
My first load pass (200 submissions, 50 judges, ~1,000 ballots) found nothing alarming. I wrote "fine at hackathon scale." Then I seeded 1,200 submissions, 520 judges, 3,600 ballots, 54,600 votes and 110,000 chained audit entries, and checked in both the seeding and measuring scripts. Three real bottlenecks:
-
The connection pool was the ceiling, not any query.
create_engine()had nopool_size, so SQLAlchemy's default of 5 + 10 = 15 connections applied to the whole process, while Starlette runs sync routes on a 40-thread pool. Request 16 was not CPU-starved; it was queued behind 14 holding connections. That is multi-second tail latency with zero errors, invisible below real concurrency, which is why pass one missed it.pool_size=20, max_overflow=20took the public gallery's p99 from 4.8s to 1.2s. -
verify_chain()hydrated 110,000 ORM objects to read twelve plain fields each. A column-only select withyield_percut verification from about 13s to about 8s. The rest is 110,000 SHA-256 calls, which is what a hash chain costs. - The vote tally built a full ORM set on every poll to produce a filter that narrowed nothing for most events. Skipping it cut concurrent p50 from roughly 11-13s to 4-7s.
Caveats I printed rather than rounded away: one laptop, and identical code varied about 2x between runs, so I report ranges. I did not cache the tally. Caching would be faster but would break the "computed on every request, never stale" property results depend on, which is a product decision, not a perf tweak.
What I cut, and what I would redo
Cut and do not regret:
- Expertise matching in assignment. Z-scoring assumes judges see comparable slices of quality, so assignment is balanced, shuffled and conflict-free. Smarter matching would have made assignment better and normalization wrong.
- Breaking ties. Equal scores share a rank (1, 1, 3). Deciding first place on a float tie-break is worse than reporting a tie.
- A job scheduler. So webhooks are not retried: five failures disable a hook. A queue is a dependency this stack does not need yet.
Would redo:
- A judge-to-tracks table. The organizers' file gives nine judges two tracks each; my judge record holds one or all, so they load as all-track judges. No loaded ballot changes, but a fresh assignment would treat them as unrestricted.
-
Cached, invalidated results.
GET /resultstakes about 1.7s at 3,600 ballots because I compute it on every request. - Secret ballots. Organizers can see who voted for what. Fixing that makes one-vote-per-person unenforceable without blind signatures, so I named it as an open problem instead of faking it.
The repo has a section called "What it does not do yet," and in places it is longer than the feature list. I think that is the most useful part of any hackathon project: not what works, but exactly where it stops, and which of your own claims survived when you went back and tested them.
Thanks to Hackathon Raptors (@partnerships_raptors) for a brief worth building. This is my entry for the Write Up Quest.
Top comments (1)
Testing your own docs against the organizers' data and publishing "it was mostly wrong" is the best part of this. The two-ballot finding is the real one: with population sd, two ballots always become +1 and -1, so those judges hand out maximum-size pushes no matter how much they actually separated the projects.
On Finding 4, there's a way to stop k being hand-picked, and a reason the k = 10 result may matter less than it looks. With 126 ballots over 41 projects, each project has about three ballots, so its normalized score has a wide spread of its own. Resample ballots (bootstrap within each project, keeping judge assignments) and recompute the ranking a few thousand times: you get a rank interval per project. If third place's interval already spans roughly 2nd to 8th at k = 3, then "drops out of the top five at k = 10" is a move inside its own noise, and the honest leaderboard shows tied bands rather than a strict order. For k itself, pick the value that best predicts held-out ballots: drop one ballot at a time, normalize with the rest, and see which k predicts the dropped z-score best. That's a data-driven version of your plateau, and it costs a few lines on top of normalization_sensitivity.py.