The sample event Hackathon Raptors gave us for Dogfood 2026 has 41 projects, 30 judges and 126 finished reviews. When my judging engine split the judges' disagreement into "this judge is just more generous" and everything else, generosity came out at about 2%, and that number shaped most of what I built and cut.
Dogfood's brief was the same for every team: in 72 hours, build the submission and judging platform Raptors will run their own events on. I entered solo. A fleet of AI coding agents (Claude Code) did the building, and I made the calls. The code, the tests and the maths are in the repo.
An engine that barely moves the ranking
The model treats every review as the project's real level, plus the judge's leniency (their tilt), plus noise. Noise means everything a correction can't fix, like one judge loving a project another found dull.
Leniency is shrunk towards zero until the data supports it, and how much data that takes is re-estimated from the event's own scores on every run. On this event a judge needs about 43 reviews before the engine trusts half of their apparent tilt. Each judge here saw about four projects. So no judge's correction is more than about 0.06 points on a 1 to 5 scale.
That looks like the engine doing nothing, but on this data there's little tilt to correct. Judges with no tilt at all, scoring the same pairs with this event's noise, still land about 0.37 apart by luck. The real judges are 0.42 apart. A test plants a real tilt, three judges made 0.6 more generous: the spread rises to 0.52 and the engine brings it back to 0.38, so it does correct a tilt that is really there.
A permutation test on the same scores finds no project differences beyond chance: in 1,529 of 2,000 shuffles, the shuffled scores spread at least as far as the real ones. On this event the engine's job is to bound what tilt could have done, and disclose it.
The usual fix, per-judge z-scores, measures each judge against the whole panel and treats all of their spread as bias. In a planning simulation before kickoff (planning was allowed, code wasn't), a test gave the harshest judge the ten strongest projects. A shrunk z-score read that judge as lenient: +0.38 against a true -0.82. My model compares each review with the other reviews of the same project, and got the sign right: -0.47.
The flat judge
One judge gave all three of their projects a 4 on every criterion. Leaving that judge out moves 19 of 41 projects, one of them (Small Relay) from 14th to 30th. The engine does leave a judge like that out, but never quietly. In simulation the same rule also flags an honest judge in 6 to 8% of panels, so it's a visible flag with its reason, and the organizer can undo it.
Two features I cut, and don't regret
The first was a chance of winning. It's the obvious feature: "Project X: 63% chance of first place." On simulated events where no project was better than any other, the method named a 50%+ favourite in 46 of 80 tracks. A percentage on a results page reads as fact, and here it would have been inventing favourites. The portal shows no chance of being first anywhere, on any page or in any API answer.
The second was letting the engine pick each pairwise question. In pairwise mode a judge just says which of two projects is better, or that they're too close to call. Having the engine choose the next pair from everyone's answers sounded smart, and I said yes after a first simulation where it picked the right winner 3.6 points more often than the simpler method. The bar for the real test was fixed in advance at 3 points. The version that refit on every judge's answers cleared it, but left some projects compared by fewer than two judges. The version that could actually ship managed 1.4, inside the noise, so I said skip it, and it was never built.
What replaced the chance of winning
Instead, each track gets a yes or no answer to one question: is the top too close to call? The engine redraws the ranking 4,000 times from its own uncertainty. If the leader comes first in fewer than 3,800 of them, the track is too close to call, and the judges can decide it. An exact tie always is.
On simulated events with no real differences, it named a winner in 0.13% of 4,000 track-runs. In the simulations where projects do differ, the winner it named was the true best 99.2% of the time, across 7,313 named winners. The downside is that in those runs it names a winner in only 20 to 63% of track-runs, so there's a lot of honest "too close to call".
The check that passed a broken engine
Every score carries a ±, and the ± has to be honest: when the engine says it's at least 95% sure, it should be right at least 95% of the time. The first check pooled all of those claims and counted.
A deliberately broken engine, with its ± cut in half, passed: right 98.4% of the time. Easy calls dominate the pool and carry the average. Checking only the claims between 95% and 99% sure dropped the broken engine to 89.7% and failed it, while the honest engine scored 98.6%.
A fixed floor for the worst single run had the opposite problem and failed the honest engine itself, in one sealed scenario. The floor became "the current engine's own worst run, minus 0.02". Both were caught before any real result existed, because every check already had to fail on a known-bad input first.
Who can see what
The organizers' checker follows redirects, so a refusal that sends you to the login page counts as a 200 and fails. Every refusal here is a real 401 or 403 from one server-side check, and a judge who asks for another judge's scores gets a 403, never their own scores instead.
At hour 68 I signed in as a judge in one tab and as the organizer in another, and changed a score in the judge tab. It looked saved, then the server refused it. A browser keeps one sign-in for all its tabs, so the judge tab had silently switched people. Now every tab notices when the sign-in changes and stops saving.
Every change goes into a hash-chained audit log in the same transaction as the change, and so does every 403 to a signed-in person. Each ballot row also gets its own random salt, because 40 projects and 3 picks make only about 60,000 possible ballots and a plain hash chain would give the votes away. A test runs that search and finds nothing.
How it was built
One Claude Code session (Opus 5.5) planned and merged, and agents built the features in parallel, each in its own git worktree. Nothing reached main until the organizers' checker (7/7) and our own hand checks (35/35) passed in a clean container. The final suite had 1,670 passing tests.
During the night run, three permission prompts reached my screen within seven minutes. I was lucky I wasn't asleep yet, since an unanswered prompt stalls the run.
So before going to sleep I set up an alarm, a watchdog outside the session that goes off when the run stalls. At 4:30 am it actually woke me up. An agent had been stuck for 16 minutes, waiting for permission to delete its own scratch files. Since then, nothing gets deleted on a night run. The mess waits for the morning.
The organizers' checker also makes one 10-second request per check, with no retry. On my loaded laptop the first gallery render took longer, and tier 1 failed once. The portal now warms its own gallery before it says it's ready.
What it doesn't do
Two judges trading favourable scores look like ordinary disagreement to the model; nothing flags that. The audit log's database triggers stop the application, not someone holding the database file, so the chain is only tamper-evident against a hash kept outside the portal. And in pairwise mode, a judge who answers "too close to call" whenever their favourite would lose is flagged in only 14 of 120 simulated panels.
Code, tests and the full maths: github.com/LippInc/dogfood-2026 (JUDGING.md has the method and its validation)
Live demo: dogfood-portal-demo.onrender.com (free instance, give it a minute to wake)
This is my Write Up Quest entry for Dogfood 2026, run by Hackathon Raptors (@partnerships_raptors).
Top comments (0)