Shipshape, the hackathon portal I built for DOGFOOD 2026, gives every team a history panel on its team page and submission page: who joined, who edited the project, when it went in. Once judging started, that panel would also have shown the team lines like these (the real log format; the names come from the test I wrote afterwards):
Jules Judge submitted a score for "Quiet Hours"
Jules Judge declared a conflict with "Quiet Hours": I mentored them
Only that team's members can open those pages. So the people being judged could read who was reviewing them, when each judge scored, whether a judge changed a submitted score, and a judge's private reason for stepping aside (its first 120 characters). Role isolation was the part of the portal I had tested hardest. None of those tests opened a team page, and none of the seven checks in the organizers' acceptance checker does either.
Teams submit projects, judges score them on a rubric, the scores are normalized so one harsh judge cannot sink a project, and the public can vote. It is Django 5.2 on SQLite. The brief came in four tiers (T1 teams and submissions, T2 judging, T3 community voting and comments, T4 an API, webhooks, signed records and more) under hard rules: 72 hours, and docker compose up must bring up a seeded portal with the network off. The organizers supply the acceptance checker and a fixtures.json every portal must load. The fixture has traps: 41 project records although the brief says "forty" (one a resubmission three minutes before the close), 30 judges, 126 scores, and a judge who gave everything a 4.
This is the list of things I believed while building it, grouped by where they broke rather than when.
I thought the features were the hard part
I read the brief as a list of features until I put "deadline actually stops submissions" next to the check that tests it. The checker POSTs {"title": "dogfood-late-submission-probe", "summary": "probe"} to a closed event and passes on any status from 400 to 499; the spec says it does not inspect why you refused. So a 404 would pass. So would a CSRF failure, or a complaint about the missing fields.
A portal could pass that check for the wrong reason, so the order of checks is the design, from the view's first version: signed in (401), event exists (404), JSON from this site (415, a 403 cross_origin, or 400), then the deadline (403 submissions_closed), and only then the team (403 no_team) and the fields (400).
body = read_json(request)
# The deadline comes before anything else about the payload: a late
# request is refused as late, whatever it contains.
try:
assert_accepting_edits(event)
except WindowError as exc:
if team is not None:
services.record_refusal(event, request.user, team, exc, "save the project", "API")
raise ApiError(403, exc.code, exc.message, deadline=iso_utc(event.submissions_close_at))
if team is None:
raise ApiError(403, "no_team", "Join or start a team for this event before submitting.")
The test runs with Django's CSRF enforcement switched on, to prove the 403 is the deadline and not a missing token. Every write path re-reads the event row inside its transaction and checks the server clock. The deadline instant itself counts as closed. A refusal is logged after the rollback, so its audit line survives.
The other sentence that was harder than it read was "with the network off", which for two days I read as a rule about runtime, until another machine ran my FROM python:3.12-slim build and it went looking for downloads. The base image and the wheels now live in the repo, built FROM scratch and installed --no-index. I picked bzip2 even though its archive is 36.9 MB to xz's 27.1 MB, because Docker unpacks xz by calling an external xz program wherever the Docker engine runs, but unpacks bzip2 itself, with Go's standard library. An offline Docker-in-Docker box then caught compose trying to pull the image from Docker Hub before building it; pull_policy: never fixed that.
I thought a z-score would fix harsh judges
In the reference example that came with the judging brief, a judge called J-harsh averages 1.71. Fixture judge jdg_02 averages 4.22. A 3 from the first is praise; from the second, a complaint. The usual fix is a z-score: how far a mark sits from that judge's own average, in units of that judge's own spread. A project's score is the mean of its judges' z-scores.
On the fixture, that formula breaks in two places. With exactly two scores, a plain z-score is always plus or minus 0.7071, because the gap between the marks cancels out of the formula. A judge who gave 3.0 and 3.1 speaks as loudly as one who gave 1 and 5. Seven fixture judges have exactly two counted scores, so plain z-scores would turn all 14 of their marks into full-strength votes. And jdg_07, the judge who gave everything a 4, has zero spread: a division by zero.
The fix is shrinkage. Each judge starts with kappa imaginary projects of population-typical spread, and their own marks pull them away from it as they pile up:
sigma_j^2 = ((n_j - 1) * s_j^2 + kappa * sigma_pop^2) / (n_j - 1 + kappa)
With kappa at its default of 5, the two-score judges separate by their actual gap: 0.277 for a third of a point, 0.534 for two thirds, 0.757 for a full point. That last one is above 0.707, so shrinkage can push a z-score up as well as down: when a judge's own spread is wider than the population's (0.65), shrinkage pulls it toward 0.65 and the z-score rises. And every z that jdg_07 produces comes out exactly 0, a neutral vote instead of a crash.
My first test of this used kappa 0.001 as "plain" and failed with AssertionError: 0.6214804438775124 not greater than 0.7. The test was wrong, not the engine: even a thousandth of an imaginary project pulls a nearly flat judge toward the population (0.7071 at kappa 0, 0.6215 at 0.001). The test now pins 0.7071 at kappa 0.
I thought plain z-scores were close enough
The reference picked kappa 5 with no evidence, so I measured: Spearman rank correlation with a hidden true quality (1.0 means the same order), 1,000 simulated events per row, seed 2026, three reviews per project. Condensed from python src/judging/engine.py --trials 1000, which gives the same numbers when rerun:
raw plain z kappa 5 kappa 25 kappa 5 beats raw
reference judges, 12 projects 0.766 0.838 0.841 0.841 75% of events
reference judges, 40 projects 0.809 0.891 0.890 0.889 98%
fixture-like, 30 judges 0.802 0.812 0.830 0.832 69%
The "plain z" column is computed at kappa 0.01, so a judge with no spread does not divide by zero; kappa 0 gives the same column to three decimals.
With the reference's judges, most of the win is just centring each judge on their own mean, and any kappa from 2 to 25 lands within about 0.003 of plain z. The third row is shaped like the real fixture, 30 judges with about four whole-number scores each, and there plain z barely moves: 0.802 to 0.812. Rerun at true plain z (kappa 0), it beats the raw average in only 55% of those events (56% at the kappa 0.01 the table uses), close to a coin flip, where kappa 5 wins 69%. So plain z-scores went.
Kappa 5 still fails to beat the raw average in about one fixture-like event in three, so the results page counts informative reviews (a review counts only if its judge has two or more scores with some spread) and flags projects ranked on fewer than two. Three fixture judges carry no signal, 5 of the 122 counted reviews, and one project is flagged: Small Relay. Normalization moves 33 of 40 projects, median 2 places. Dry Harbour climbs from 29th to 8th: the judge who dragged it down scored nothing else, so that lone mark normalizes to neutral, and most of its other judges marked it above their own averages (jdg_26 gave 4.67 against a 3.70 average).
I thought my rankings were deterministic
JUDGING.md once promised that ties were broken by raw average, then project id, "so the order is always deterministic". It quoted the fixture's biggest moves from my Windows machine: Glass Beacon up 10 places, Paper Anchor down 10. Before handing judging over, I rechecked every documented figure against a fresh seed, and that was the first time the proof ran inside a Linux container. It said up 8 and down 9.
Three projects average exactly 3.50, and the old code handed out positions, not ranks, so the tie went to whichever float came out a hair larger:
# before: positions 1, 2, 3, even for equal averages
for position, row in _rank(reviewed, key=lambda r: (-r.raw_mean, -r.z_mean, r.project.pk)):
row.raw_rank = position
# after: competition ranking (1, 2, 2, 4), values compared at 9 decimal places
ordered = sorted(rows, key=lambda r: -round(value(r), 9))
| Raw average 3.50 | Raw rank, Windows | Raw rank, container | Raw rank, fixed | Normalized rank |
|---|---|---|---|---|
| Glass Beacon | 21 | 19 | 19 | 11 |
| Open Kiln | 19 | 20 | 19 | 21 |
| Paper Anchor | 20 | 21 | 19 | 30 |
The fix was competition ranking at nine decimals, a fixed score read order and test_equal_scores_share_a_rank_whatever_order_they_were_summed_in, which took the suite from 125 tests to 126. The rebuilt container and a fresh local seed then printed the proof identically.
My tests could not see it because they all ran on one machine. Working the arithmetic again, I think my docs also blame the wrong thing: they say "summation order" between two databases, but the weighted sum always walks the criteria in a fixed order. The likely culprit is Python:
Python 3.11, each criterion weighted 1/3
marks (3, 4, 4) sum to 3.666666666666666
marks (3, 3, 5) sum to 3.6666666666666665
Glass Beacon's raw mean: 3.4999999999999996
The Windows venv runs 3.11, so Glass Beacon sorts last of the three despite the best z-score. The container runs 3.12.14, whose sum() switched to a more accurate algorithm: all three are exactly 3.5, the old key falls through to the z-score, and Glass Beacon comes first. Both orders I saw follow from that.
A day later another test failed only in the Linux image, with AssertionError: 3.0 == 3.0. It compared whichever of two tied projects came first, both at 3.0 before and after the change it checked, so it had never tested what it claimed.
I thought quadratic voting would stop a loud minority
The community vote was going to be quadratic, the textbook answer to a loud minority: n votes on one project cost n squared credits. My first simulation disagreed. A fifth of the voters backing one weak project won 100% of events, worse than one person, one vote at 96%, because a bloc member puts 5 votes on the target while a sincere voter's favourite gets 2 or 3. So I added a ceiling of 3 votes per project per ballot and reran with 1,000 events per row, 40 projects and 300 voters. "Seen" is how many projects a voter actually looks at:
| Method | Best project wins, nobody gaming, 10 seen | Same, 20 seen | 20% bloc wins, 10 seen | 10% bloc, 3 identities each, 10 seen |
|---|---|---|---|---|
| One person, one vote | 79% | 90% | 95% | 100% |
| Quadratic, no ceiling | 64% | 84% | 100% | 100% |
| Quadratic, ceiling of 3 | 48% | 82% | 57% | 100% |
At 20 projects seen, the ceiling held a 20% bloc to 0% wins where the uncapped version let it win 40%. The cost is in the first column: when nobody games the vote, one person, one vote crowns the best project (the highest hidden quality) 79% of the time at 10 seen; my capped method manages 48%. It orders the whole field better (rank correlation 0.984 against 0.926), but the winner is the number people remember.
The last column is the same for every method: a bloc whose members each hold three identities wins every event. Past that point the work is in who gets a ballot: three access modes (an open link, email links stored only as hashes, or signed-in accounts) and a detector that flags bursts, shared devices and aliased addresses.
Ballot order mattered too. With one shared order and voters looking at about 10 projects, 94% of the top three came from the first ten listed, and the best project won 13% of the time. A stable shuffle per voter, seeded by an HMAC of who the voter is and never stored, lifted that to 51%.
I thought isolation meant guarding the judges' endpoints
By the time the T3 brief repeated the isolation requirement, judge queries started from the signed-in judge, anything outside a judge's assignments answered 404 like a missing id, and a request for a peer's scores was refused with 403 before any lookup, tested with three spellings of another judge and one id that belongs to nobody.
This time the question was not what a judge can reach, but where judging data ends up. A grep for history in the templates, then a list of judging log calls that pass team=, found it. In T1, the team history panel read Activity.objects.filter(event=event, team=team) with no filter on the kind of entry. In T2, five judging actions (assigning by hand, unassigning, submitting a score, changing one, declaring a conflict) started tagging their audit lines with the team. Neither change was wrong alone, but together they sent T2's judging lines into a panel T1 had built to show teammates who joined and who edited.
Every isolation test probed judging endpoints, and so do the checker's; this leak ran from judges to the team through a generic audit log. The seed imports the fixture's 126 scores without logging anything, which is likely why walkthroughs on seeded data showed nothing odd. The fix is one helper both pages must go through:
# Entries a team must never see in its own history: who judges their project,
# when a judge scored it, a judge's conflict and their reason, and anything
# about the community vote. They stay in the organizers' activity log.
STAFF_ONLY_VERBS = ("score.", "judge.", "judging.", "voting.")
def team_history(event, team, limit):
entries = Activity.objects.filter(event=event, team=team)
for prefix in STAFF_ONLY_VERBS:
entries = entries.exclude(verb__startswith=prefix)
return entries.select_related("actor")[:limit]
Its test plants a conflict reason in capitals, I MENTORED THEM, and asserts it reaches neither team page but still reaches the organizers' log. The weakness was built in from the start: a deny list keyed on how verbs are spelled.
"Hidden looks exactly like missing" failed once more the next day, in a threat-model test comparing the API's answer for a draft with its answer for an id that does not exist:
AssertionError: {'error': 'not_found', 'detail': 'No such project, or nothing you can see.'} != {'error': 'not_found', 'detail': 'Nothing here, or nothing you can see.'}
Two polite 404s, one sentence apart: enough for anyone counting up through project ids to tell a hidden draft from an empty slot. The hidden case now raises the same Http404 as a missing one.
Some of my tests pass because the attack still works
When the tests moved out of the Django project, running the suite from inside src/ printed Ran 0 tests in 0.000s OK: a pass that touched nothing. The test runner now defaults to the tests/ folder, so it cannot quietly find zero tests.
The threat model, one of the two bonus challenges I took on, lists 43 attacks: 28 stopped, 3 capped or flagged, 1 logged only and 11 open. The doc and the tests share attack ids, and four of the open attacks have tests with _gap_ in the name that pass because the attack works:
def test_SY5_gap_a_fresh_browser_gets_a_fresh_ballot_on_an_open_link(self):
url = reverse("voting:ballot_link", args=[self.event.slug, self.config.link_token])
for _ in range(2):
browser = Client() # a private window, or cookies cleared
browser.get(url)
browser.post(reverse("voting:api_ballot", args=[self.event.slug]) + f"?link={self.config.link_token}",
data=json.dumps({"votes": {str(self.other.pk): 1}}), content_type="application/json")
self.assertEqual(Ballot.objects.filter(event=self.event).count(), 2)
Close that hole and the test goes red, which forces a change to the doc that admits it. A suite that records only what works cannot tell you when your list of weaknesses has gone stale.
Writing the threat model turned up four more holes besides the 404 wording. A passed deadline could be moved later under scores already given (now refused once a score or ballot exists). Submitted projects were public before the deadline, handing early ideas to teams still working (now private by default). One inbox could cast several ballots (more on that below). And nothing looked for a judge boosting a friend.
For that last one I flag a project when one judge's normalized score sits far from the rest of its panel. I first picked a threshold of 2.0, because only 4 of 106 honest gaps on the fixture reached it. Then the simulation ran: at 2.0 the flag caught only 25% of single-judge boosts. The sweep, over 1,000 simulated events:
threshold 1 colluder caught 2 colluders caught honest projects flagged
1.0 86% 84% 29.7%
1.25 75% 70% 14.5%
1.5 58% 56% 5.8%
1.75 40% 40% 2.1%
2.0 25% 28% 0.6%
I settled on 1.5: it catches 58% of single boosts and flags 5.8% of honest projects, two or three in every 40, a list short enough for an organizer to read. On the real fixture the flag fires far more often: on 8 of the 32 projects it can check, against the simulated 5.8%, and the 8 projects with only two reviews can never be checked. Two colluders on a three-judge panel still get a bottom-half project into the top three in 24% of events. The cheapest defence was already there: random assignment puts a given judge on a given friend's panel about one time in ten.
On its first run, my HTTP attack probe reported two attacks as working. Both were bugs in the probe, and the doc still lists them. Once fixed, the probe ran 14 attacks and the portal stopped all 14.
I thought an email address was a voter
Until 29 hours into the build (for the first 20 hours that email ballots existed), an email ballot belonged to the address as typed, under a unique constraint on (event, email). My own anti-abuse detector knew better: its canonical_email drops +tags and Gmail dots, so ada+2@x.org is ada@x.org. But the ballot, the per-address link limit and the own-team check all used the typed address, so one inbox could hold one ballot per spelling, merely flagged. My docs even claimed +tags were dropped everywhere.
The fix touched five code paths and zero migrations. Ballot.email started holding a different kind of value under the same constraint name, one_ballot_per_email, with a compatibility clause in place of a data migration:
if voter.kind == Ballot.Kind.EMAIL:
# Ballots from before inboxes were used keep the address as typed.
return qs.filter(email__in={inbox(voter.email), voter.email}).order_by("pk").first()
Replaying the change shows what that shortcut costs:
- A ballot saved before the change as
kim+1@x.orgis missed when the inbox returns askim+2@x.org, so one inbox gets two ballots. Only databases that held email ballots before the change are exposed; a fresh seed has none. - The archive importer writes addresses raw, so
lee+1@x.organdlee+2@x.orgimport as two ballots. - With no stored inbox, the own-team check runs a leading-wildcard
LIKEand filters in Python. With 50 Gmail accounts in the table, all 50 were loaded for one check. - The typed spelling is lost from the ballot (
Mal.Lory@googlemail.comis stored asmallory@gmail.com; only the email-link row keeps it).
What I would write now: keep email_as_typed, add an indexed inbox column computed at write time, put the unique constraint on (event, inbox), and give EmailPass and User the same column. The importer could not get around a rule the database enforces. The RunPython backfill would fail on any inbox already holding two ballots, and that failure is useful: it forces the "which ballot wins" decision my compatibility clause skips. One cost stays: on a provider where a+b@ and a@ are different people, they share a ballot.
The second redo is the audit log. Activity serves three audiences (the team, organizers, platform admins), and who may read a row is decided by its verb's spelling: of 25 verbs logged with a team, 5 are hidden only by their prefix. I would add an audience column, set when the line is written and defaulting to staff, so a verb nobody thought about stays hidden instead of showing up on a team page.
One choice I would keep: nothing derived is stored. The whole fixture ranking is recomputed per request in about 25 to 30 ms and 6 queries on my machine. So when I noticed, about 32 hours in, that organizers had no button to publish the judges' results, the fix was three columns on the judging settings (results_published_at, results_published_by, share_feedback) and no stored ranking to freeze: the public page runs the same calculation as the organizers' one.
One rubric per event, on purpose
The reference design allowed a rubric per track. Shipshape has one per event, and I would make that cut again. A z-score compares a judge's mark with that judge's own mean and spread, which only means something on one scale. In the fixture, 7 of the 30 judges scored in two tracks, giving 50 of the 126 scores, including jdg_24 and jdg_26, the two judges whose session cookies .dogfood.toml hands the checker. Per-track rubrics would have put 40% of the scores on two scales inside one judge's z-score. Organizers keep weights that stay editable after scoring starts, which is cheap because totals are recomputed on every request.
I also skipped reliability weighting of judges, because early scores would decide how much later ones count. I did not attempt the pairwise judging bonus, and I will not dress that up as a principle: I took on the threat model and API-first instead.
I thought I should claim every tier I built
The checker has seven checks, three for T1 and four for T2, and its rule for a verified tier fits in three lines (plus a loop that drops any tier above one that failed):
verified = [t for t in TIERS
if any(c.tier == t for c in checks)
and all(c.ok for c in checks if c.tier == t)]
There are no T3 or T4 checks, so any is false for them and no portal can get them verified. The spec says overclaiming is the one thing that costs points. I built all four tiers, and for about a day and a half my report ended claimed T1 T2 T3 T4, verified T1 T2; my first commit carried it. Then I split the claims: .dogfood.toml claims what a program can confirm, and the README claims T3, T4 and the two bonuses next to the tests behind them. All seven checks pass, and the report now ends:
T2 judge cannot see peer scores ...... PASS
T2 participant blocked ............... PASS
T2 csv export works .................. PASS
claimed T1 T2, verified T1 T2
I held the API docs to the same rule. Checking every /api/v1/ answer against its OpenAPI schema during the tests found five places the docs were wrong, and showed my 307 tests reached only 73 of 134 endpoints, so a run now fails if any endpoint goes unexercised. A Schemathesis run reached about 1,000 requests with no server error before I stopped it, and the docs say that rather than claiming a pass.
Where the failures actually were
The formulas were the part I could prove fastest, because a simulation has no stake in my being right. The method I shipped beats the raw average in 69% of fixture-like events; plain z-scores do it in 55%. I would rather print both numbers than print a ranking and call it fair. Every serious failure here happened somewhere else: the wording of a 404, which Python summed the marks, which column a ballot was keyed on, where an audit line ended up.
Checking facts for this post, I seeded a fresh fixture, set the results date three days ahead, gave "Best in show" to Glass Signal and opened the winning team's page as one of its members. Awards stay private until the results date, so organizers can decide winners before the ceremony, and the public awards list correctly returned 0 rows. The team page returned 200 with "Best in show" in its history, from this line:
log_activity(event, actor, "award.given", f"gave “{prize.name}” to “{project.title}”",
team=project.team, project=project)
award. is not on the deny list. No test covers it, and it is still there as I write this: the opening scene again, with a different verb. Adding award. to the list would close this one. The audience column would have kept it closed without anyone having to remember.
Try it
git clone https://github.com/manusingh090/Shipshape.git
cd Shipshape
docker compose up
No network is needed. Open http://localhost:8080, choose Sign in and press one of the Be buttons: Be Rosa is the organizer, Be Priya a team captain, Be Tom a participant with no team, Be Ada and Be Diego are judges. The tables come from python src/judging/engine.py --trials 1000 (add --collusion for the collusion simulation) and python src/voting/method.py --trials 1000 (add --seen 20 for the 20-seen column and --order for ballot order), which need only plain Python.
References
- The DOGFOOD 2026 brief and acceptance checker, in the repo as
tests/acceptance/spec.mdandtests/acceptance/run.py - JUDGING.md in the repo: section 4 (normalization and the Monte Carlo proof), section 10 (the vote simulations), section 11 (the threat model)
- Schemathesis: https://schemathesis.readthedocs.io/
- Shipshape: https://github.com/manusingh090/Shipshape
- Google Slides : https://docs.google.com/presentation/d/1z8AnsudKyrrzO9azSSoI51wdYGTVh7hNFrpsKxTI_Eg/edit?usp=sharing
Top comments (1)
Just a wow ! Project