DEV Community

realist
realist

Posted on

Three Judges Scored the Same Project Twice. They Disagreed With Themselves by 1.11 Points.

Why Evenhand ranks 40 hackathon projects in 2 groups, and shows its receipts.

A hackathon ranking is forty claims that each project beat the one below it. We wanted to know how many of those claims the scores could actually support.

The organisers' sample data for DOGFOOD 2026 hides one project twice. Dry Harbour is filed as prj_07 and again as prj_41: one team, one project, submitted twice. It is a trap. A judging portal is supposed to notice and not count the project twice.

Three of the sample data's invented judges scored both copies. Each score is the mean of three criteria on a 1–5 scale.

Judge First copy (prj_07) Second copy (prj_41) Gap
jdg_19 2.33 3.67 +1.33
jdg_21 3.67 4.67 +1.00
jdg_26 4.67 3.67 −1.00

Same project, same judge: 1.11 points apart on average. Our own model of judging noise says two honest looks at the same work should differ by about 0.68. Two judges went up and one went down, so this is not the second copy simply looking better.

Evenhand, the portal we built, was supposed to turn 121 of these scores into a ranking of 40 projects. It returned two groups.

Three made-up judges are not a study, and we will be careful about that below. This is the write-up of what that number did to everything else we built, what broke on the way, and the one decision we would take back.

What Evenhand is, and who built it

DOGFOOD 2026's premise: every team builds the same hackathon platform in 72 hours, and the winner is forked to run future Hackathon Raptors events. You build the thing that will judge you. The organisers supplied a spec, an acceptance checker (run.py, seven probes) and sample data: 41 projects, 30 judges, 8 tracks and 126 reviews, with traps planted in it.

Evenhand is our entry: a public gallery, a judge console where judges cannot see each other's scores, and rankings corrected for harsh and generous judges. One command brings it up, seeded, with the network off:

$ docker compose up
...
seeded. test logins (DEMO_MODE=true: demo only, never use in production):
  organizer    Authorization: Bearer dev-organizer-7f2a    (organizer@evenhand.local)
  judge_a      Authorization: Bearer dev-judge-a-91bc      (jdg_24 Diego Herrera)
  judge_b      Authorization: Bearer dev-judge-b-44de      (jdg_26 Jonas Vogel)
  participant  Authorization: Bearer dev-participant-2e88  (priya1@example.org)
Enter fullscreen mode Exit fullscreen mode

Under it: a NestJS API that owns every rule, a Next.js web app that is only one of its clients, Postgres, and a judging engine in plain TypeScript with no I/O, so the code that decides who wins can be tested on its own.

Each decision is a dated record in docs/decisions, 26 of them, written when the call was made rather than afterwards. The bugs below were found by the sample data or by running it offline, not by reading code.

A trap used as an instrument

You normally cannot measure how much a judge disagrees with themselves, because nobody judges the same thing twice. The duplicate made that happen by accident. In statistics this is called a test–retest, and the trap gave us three pairs of one for free.

The 0.68 comes from the model. After accounting for which project was judged and how generous each judge is, the leftover noise on a single review is about 0.60 points. Two independent draws of that noise differ by 2σ/√π ≈ 0.68 on average. The three judges differed by 1.11. They disagree with themselves more than the model expects them to disagree with anyone.

A duplicate entry is a free test–retest. Measure it before you trust a ranking built on the same judges.

Two caveats belong here, not in a footnote.

  • These judges are invented. The organisers generated them. The 1.11 tells you how the sample data was made, not how people judge. We make no claim about human nature.
  • Three pairs is three pairs. We show all three gaps above and print no p-value, because an interval on n = 3 would be decoration.

This is not a DOGFOOD problem. Here each project has two to five reviews, so a one-point wobble per judge is enough to swap neighbours anywhere in the ranking. And no event needs a lucky duplicate to measure that wobble. Show two or three judges one calibration entry twice, under two names, and you learn how far a judge moves on identical work before you publish anything that depends on it.

The claim we keep is narrow: on this data, there is not enough signal to put 40 projects in order, and the software should notice that by itself instead of a human spotting it by hand. Evenhand does. Every number in this section comes from one command, which regenerates our Normalization Proof from the sample data:

git clone https://github.com/Len3hq/evenhand && cd evenhand
npm install
npm run proof -w @evenhand/judging-engine   # writes docs/proof/normalization.md
Enter fullscreen mode Exit fullscreen mode

Section 7 of that file holds the three pairs and section 4 the two groups.

What broke, in the order it hurt

Four things that looked right until the sample data, or a start with the network off, showed otherwise.

The textbook fix deletes four judges

The usual correction for harsh and generous judges is the per-judge z-score:

// the textbook correction (we did not ship it)
const z = (score - judgeMean) / judgeSd;
Enter fullscreen mode Exit fullscreen mode

On this data judgeSd is zero four times. jdg_07 gave 4/4/4 to every project. jdg_12 and jdg_23 have one review each. jdg_19's three reviews happen to share an average. Divide by zero and those four judges quietly leave the ranking. The formula also assumes every judge saw projects of average quality, which is false when judges are matched to tracks and some tracks are stronger.

So we estimate everything at once. Each score is the overall average, plus how good the project is, plus how generous the judge is, plus noise:

score(project p, judge j) = μ + q_p + b_j + ε
Enter fullscreen mode Exit fullscreen mode

Every project quality and every judge's generosity is pulled towards zero unless the data argues otherwise. A judge with one review barely moves. A flat judge's generosity is absorbed without dividing by anything. A project with two reviews is pulled harder towards the middle and shown with a wider error bar.

For the maths people. This is ridge regression, the approach of Roos, Rothe and Scheuermann (AAAI 2011), solved exactly with a hand-written Cholesky factorisation in the pure engine. The two penalties are chosen by exact leave-one-out error from the hat matrix, e_i / (1 − h_ii), which a test checks against real refits. We planned a grid of 0.25 to 8 and had to widen it to 0.125 to 64, because the choice hit both edges. Rows are summed in a fixed order, so the result is identical to the bit whatever order reviews arrive in.

Before you trust a correction, count whom it quietly drops.

Three different marks, one average

The progress dashboard warns organisers about a judge who gives every project the same marks. Our first version compared each judge's weighted scores. The sample data showed why that is wrong: it flagged jdg_19 as well, whose marks were 3/5/3, 3/4/4 and 5/4/2. All three average 3.67 by chance. That judge told the projects apart.

A test built around jdg_07 alone would have passed; it took the full sample data to show the extra name. The fix compares marks criterion by criterion, and the case that fooled us is now a test:

it('does not flag different marks that happen to average the same (fixture jdg_19)', () => {
  expect(
    isFlat([
      { functionality: 3, quality: 5, innovation: 3 },
      { functionality: 3, quality: 4, innovation: 4 },
      { functionality: 5, quality: 4, innovation: 2 },
    ]),
  ).toBe(false);
});
Enter fullscreen mode Exit fullscreen mode

Only jdg_07 is flagged now. Test a warning against the case it must not fire on, not only the case it must.

The merge that would have deleted a judge

The obvious rule for Dry Harbour is "keep the newer copy, drop the older one and its reviews". Dropping prj_07's reviews deletes jdg_01's only score in the whole file. "The later score wins" cannot be computed either, because the sample scores carry no timestamps.

So confirming a duplicate moves reviews that exist only on the older copy across to the newer one. Where a judge reviewed both, the newer review counts and the older one is set aside, never deleted. prj_41 ends with 6 counted reviews, and undoing the decision restores every row; a test compares them with a snapshot taken before.

The schema had its own version of this fight. "One live submission per team" is a database index, and "both copies stay visible until an organiser decides" breaks it on import. The answer was one more column:

CREATE UNIQUE INDEX "submissions_one_live_per_team"
  ON "submissions" ("event_id", "team_id")
  WHERE "superseded_by_id" IS NULL AND NOT "duplicate_hold";
Enter fullscreen mode Exit fullscreen mode

Before a cleanup rule deletes anything, count whose only record it deletes.

The server that wanted the internet

The rule was that docker compose up must work with no network. Prisma picks its database engine binary by the OpenSSL version it detects at install time. When a build stage saw a different OpenSSL from the runtime image, the container tried to download a matching engine on start, which fails with no network.

A developer's machine always has a network, so this bug cannot appear there. It only appears on the machine the rule is about. Every stage now installs the same OpenSSL, and the build refuses to finish rather than letting the start fail later:

# Fail the build (not the offline start) if the schema engine for this OpenSSL is missing.
RUN ls node_modules/@prisma/engines/schema-engine-*openssl-3.0.x >/dev/null \
 || { echo "Prisma schema engine for OpenSSL 3 missing: runtime would try to download it" >&2; exit 1; }
Enter fullscreen mode Exit fullscreen mode

If a step needs the network, make the build fail, not the start.

The number we did not want to publish

Here is the number that makes our model look bad. Hide one review, predict it from all the others, and repeat for every review. Guessing "the average of all the other scores" gives a mean squared error of 0.3996. The full model, which knows the project and the judge, gives 0.3961. That is 0.9% better than guessing the average.

On this data, knowing which project a score belongs to barely helps you guess the score. The penalties the data picks, λ_b = 32 and λ_q = 16, mean "shrink hard": every project is pulled most of the way back to the middle. Halving or doubling them barely moves the order (Spearman ρ 0.999, the same top five), so no amount of tuning was going to make the ranking decisive.

We could have left the number in a notebook and printed 40 confident ranks. Instead the results page prints it, as the signal level, alongside the ranking it undermines, and asks readers to read tie groups rather than single ranks when it is low.

Print how much better your model does than the plain average, next to the ranking it produces.

The part we are proudest of is a refusal

Evenhand will not tell you who came 17th.

A ranking run still produces an ordered list. Each project gets its raw mean, its corrected score with an error bar, a rank, and a one-line reason. It also gets a tie group: projects within one pooled standard deviation of the group's leader are marked as equal, and the page says to read them that way. On the sample data that gives 16 projects in the first group and 24 in the second. The corrected scores of all 40 fit between 3.46 and 3.72, with an error bar of about ±0.15 on each. Printing them as 1 to 40 would be claiming precision the data does not have.

A reason reads like this, so an organiser can see why a project moved:

+0.12: its judges were on average 0.12 less generous than the rest; pulled towards the mean (2 reviews)
Enter fullscreen mode Exit fullscreen mode

chart

Each hollow dot is a project's raw average; the filled dot and bar are its corrected score ± 1 SD. Most of the raw spread is noise the model cannot tie to the project, so it pulls every score towards the middle. Judges' generosity moves no score by more than 0.1 points.

The refusal has teeth in three places:

  • Every run is a receipt. It stores a SHA-256 hash of its inputs and of its result. Results stay hidden from everyone until an organiser publishes a run.
  • A stale run cannot be published. If a review, a weight or a duplicate decision changes after the run, the input hash no longer matches and publishing is refused until someone runs it again.
  • The proof cannot drift from the code. The Normalization Proof is generated by code, and a test regenerates it and fails if the committed copy differs by a single byte. The numbers in this post are the numbers the code computes today.

The obvious question is how anyone awards first, second and third from two groups. The list is still ordered, so the order is there for anyone who wants it. But the decision lives in the top group: 16 projects the data cannot separate. Two honest ways through are prizes per track, which compare judges who saw similar work, or a final human panel on the top group. A project with a single final review is listed and never ranked at all.

A third decimal place is not a tiebreak.

Two other refusals

The server refuses before it looks. When judge B asks for judge A's scores, the API refuses before asking the database whether judge A exists. A real judge and a made-up one get byte-identical 403s, so nobody can probe for valid ids. A test tries every role against every operation on every change: 448 expected status codes.

it('refuses judge_b before looking anything up: same answer for a real and a fake judge', async () => {
  const real = await t.http().get('/api/judges/jdg_24/scores').set(headers.judge_b);
  const fake = await t.http().get('/api/judges/jdg_nope/scores').set(headers.judge_b);
  expect(real.status).toBe(403);
  expect(fake.body).toEqual(real.body);
});
Enter fullscreen mode Exit fullscreen mode

Decide permission from who is asking, before you look up what they asked for.

The stack refuses the internet. The database and API sit on a Docker network with no route out, so "no external calls at runtime" holds even if someone adds code that tries. npm run drill proves it: it clones a fresh copy, builds it, cuts the network, checks that a request to the internet now fails from both containers, and then runs the organisers' checker (7/7) and 46 browser steps inside the sealed network.

Prove an offline rule by cutting the network, not by reading the code.

What it does not do

A post about refusing to overclaim should list its own gaps. These are in the README too, written so a reviewer does not have to find them.

Limit What it means
Our best accuracy numbers are on data we generated On 20 seeded synthetic events the model recovers the true order better than raw means or z-scores (Spearman ρ 0.92 against 0.90 and 0.74). That proves the code finds a planted truth. It does not prove real judges behave like our generator.
A judge's scale is not modelled One judge who uses 1–5 and another who only uses 3–4 look the same to the model. With 1 to 11 reviews per judge there is too little data to estimate it.
The audit log stops at the database superuser Triggers refuse UPDATE, DELETE and TRUNCATE on the log, but a superuser can switch triggers off. Hash-chaining is not built.
Per-IP rate limits can be dodged Behind the bundled proxy a client can rotate X-Forwarded-For. A reverse proxy in front has to overwrite it.
No email and no TLS in the box Invite links are shared by hand; put a TLS-terminating proxy in front for anything beyond a laptop.
Not built Image uploads and community voting.
Receipts are hashes, not signatures The SHA-256 hashes prove a published result matches its inputs. They do not prove who published it.

The numbers

Number What it counts
1.11 average self-disagreement of three sample judges on one project, 1–5 scale
0.68 the gap the model's noise predicts for the same comparison
0.9% how much better the model predicts a hidden review than the plain average
2 tie groups for 40 projects
121 reviews in the ranking (126 in the file; the older duplicate's 5 held out)
448 expected status codes in the role-isolation test
7 / 7 organisers' checker probes passing, also with the network cut
46 / 46 browser steps passing, with no Content-Security-Policy violation
107 + 665 unit and end-to-end tests
26 decision records, written as the decisions were made
43 hours from first commit to last

The decision we would take back

The hours we spent making the model exact: the hand-written Cholesky solve, exact leave-one-out, the widened grid. It is good engineering, and on this data it changed nothing a reader can see. The signal is 0.9%, and the order barely moves whatever the penalties are.

The same hours could have hash-chained the audit log, so that every entry carries a hash of the one before it and a rewritten history no longer lines up. Today, someone with superuser access to the database can switch the triggers off and rewrite the trail without leaving a mark. Our pitch is that rankings come with receipts. A portal built on that pitch should have spent its last hours making the receipts hold all the way down.

What we would tell someone building one of these

  • Find the test–retest before you rank. A duplicate, a resubmission, a calibration entry: anywhere a judge saw the same work twice tells you how much of a score is noise.
  • Print the boring baseline. Report how much better your model predicts a hidden score than the plain average does. If it is 1%, print groups, not ranks.
  • Count whom a correction drops. Z-scores, minimum-review rules and duplicate merges all remove people quietly. Count them by name.
  • Refuse before you look up. Decide permission from the caller, then test that a real and a fake id get the same answer.
  • Generate the proof from the code. Our proof document is rebuilt by a test, and the build fails if one byte drifts.
  • Write down what you did not build, before a reviewer finds it.

Try it

git clone https://github.com/Len3hq/evenhand && cd evenhand
docker compose up                                    # seeded portal on http://localhost:8080
curl -H 'Authorization: Bearer dev-judge-b-44de' \
  localhost:8080/api/judges/jdg_24/scores            # 403: not your scores
npm install && npm run proof -w @evenhand/judging-engine   # regenerates the Normalization Proof
npm run drill                                        # fresh clone, network cut, checker and browser steps
Enter fullscreen mode Exit fullscreen mode

The first build needs the internet, for about 3 to 4 minutes. After that, nothing does.

Where the hours went

We expected the hard part to be the maths. It was not. Writing a ridge regression and a Cholesky solver is the fast part now; ours was built with Claude Code. What took the hours was deciding what the output is allowed to claim: that two scores 0.03 apart are not a ranking, that a flat judge is not a broken one, and that a duplicate is a measurement before it is a cleanup job.

The organisers planted Dry Harbour to see whether a portal would count one project twice. It turned out to be the only place in the data where a judge could be measured against themselves, and that measurement set the shape of everything else we built.

If you take one thing from this: before you print a rank, find out how far one judge moves on the same work. On our data it was more than a point.

Built for DOGFOOD 2026 by Hackathon Raptors. #hackathonraptors #dogfoodhack

Top comments (0)