<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Manjeet Singh</title>
    <description>The latest articles on DEV Community by Manjeet Singh (@manusingh).</description>
    <link>https://dev.to/manusingh</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4152835%2Fa7a07ce6-6009-4212-9ba0-1f8c03285e46.png</url>
      <title>DEV Community: Manjeet Singh</title>
      <link>https://dev.to/manusingh</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/manusingh"/>
    <language>en</language>
    <item>
      <title>Seven green checks and a team page that named its judges</title>
      <dc:creator>Manjeet Singh</dc:creator>
      <pubDate>Fri, 02 Oct 2026 06:29:50 +0000</pubDate>
      <link>https://dev.to/manusingh/seven-green-checks-and-a-team-page-that-named-its-judges-57en</link>
      <guid>https://dev.to/manusingh/seven-green-checks-and-a-team-page-that-named-its-judges-57en</guid>
      <description>&lt;p&gt;Shipshape, the hackathon portal I built for DOGFOOD 2026, gives every team a history panel on its team page and submission page: who joined, who edited the project, when it went in. Once judging started, that panel would also have shown the team lines like these (the real log format; the names come from the test I wrote afterwards):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Jules Judge submitted a score for "Quiet Hours"&lt;/p&gt;

&lt;p&gt;Jules Judge declared a conflict with "Quiet Hours": I mentored them&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Only that team's members can open those pages. So the people being judged could read who was reviewing them, when each judge scored, whether a judge changed a submitted score, and a judge's private reason for stepping aside (its first 120 characters). Role isolation was the part of the portal I had tested hardest. None of those tests opened a team page, and none of the seven checks in the organizers' acceptance checker does either.&lt;/p&gt;

&lt;p&gt;Teams submit projects, judges score them on a rubric, the scores are normalized so one harsh judge cannot sink a project, and the public can vote. It is Django 5.2 on SQLite. The brief came in four tiers (T1 teams and submissions, T2 judging, T3 community voting and comments, T4 an API, webhooks, signed records and more) under hard rules: 72 hours, and &lt;code&gt;docker compose up&lt;/code&gt; must bring up a seeded portal with the network off. The organizers supply the acceptance checker and a &lt;code&gt;fixtures.json&lt;/code&gt; every portal must load. The fixture has traps: 41 project records although the brief says "forty" (one a resubmission three minutes before the close), 30 judges, 126 scores, and a judge who gave everything a 4.&lt;/p&gt;

&lt;p&gt;This is the list of things I believed while building it, grouped by where they broke rather than when.&lt;/p&gt;

&lt;h2&gt;
  
  
  I thought the features were the hard part
&lt;/h2&gt;

&lt;p&gt;I read the brief as a list of features until I put "deadline actually stops submissions" next to the check that tests it. The checker POSTs &lt;code&gt;{"title": "dogfood-late-submission-probe", "summary": "probe"}&lt;/code&gt; to a closed event and passes on any status from 400 to 499; the spec says it does not inspect why you refused. So a 404 would pass. So would a CSRF failure, or a complaint about the missing fields.&lt;/p&gt;

&lt;p&gt;A portal could pass that check for the wrong reason, so the order of checks is the design, from the view's first version: signed in (401), event exists (404), JSON from this site (415, a 403 &lt;code&gt;cross_origin&lt;/code&gt;, or 400), then the deadline (403 &lt;code&gt;submissions_closed&lt;/code&gt;), and only then the team (403 &lt;code&gt;no_team&lt;/code&gt;) and the fields (400).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;read_json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# The deadline comes before anything else about the payload: a late
# request is refused as late, whatever it contains.
&lt;/span&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;assert_accepting_edits&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;WindowError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;team&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;record_refusal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;team&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;save the project&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;API&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ApiError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;403&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;deadline&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;iso_utc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;submissions_close_at&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;team&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ApiError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;403&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no_team&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Join or start a team for this event before submitting.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The test runs with Django's CSRF enforcement switched on, to prove the 403 is the deadline and not a missing token. Every write path re-reads the event row inside its transaction and checks the server clock. The deadline instant itself counts as closed. A refusal is logged after the rollback, so its audit line survives.&lt;/p&gt;

&lt;p&gt;The other sentence that was harder than it read was "with the network off", which for two days I read as a rule about runtime, until another machine ran my &lt;code&gt;FROM python:3.12-slim&lt;/code&gt; build and it went looking for downloads. The base image and the wheels now live in the repo, built &lt;code&gt;FROM scratch&lt;/code&gt; and installed &lt;code&gt;--no-index&lt;/code&gt;. I picked bzip2 even though its archive is 36.9 MB to xz's 27.1 MB, because Docker unpacks xz by calling an external &lt;code&gt;xz&lt;/code&gt; program wherever the Docker engine runs, but unpacks bzip2 itself, with Go's standard library. An offline Docker-in-Docker box then caught compose trying to pull the image from Docker Hub before building it; &lt;code&gt;pull_policy: never&lt;/code&gt; fixed that.&lt;/p&gt;

&lt;h2&gt;
  
  
  I thought a z-score would fix harsh judges
&lt;/h2&gt;

&lt;p&gt;In the reference example that came with the judging brief, a judge called J-harsh averages 1.71. Fixture judge &lt;code&gt;jdg_02&lt;/code&gt; averages 4.22. A 3 from the first is praise; from the second, a complaint. The usual fix is a z-score: how far a mark sits from that judge's own average, in units of that judge's own spread. A project's score is the mean of its judges' z-scores.&lt;/p&gt;

&lt;p&gt;On the fixture, that formula breaks in two places. With exactly two scores, a plain z-score is always plus or minus 0.7071, because the gap between the marks cancels out of the formula. A judge who gave 3.0 and 3.1 speaks as loudly as one who gave 1 and 5. Seven fixture judges have exactly two counted scores, so plain z-scores would turn all 14 of their marks into full-strength votes. And &lt;code&gt;jdg_07&lt;/code&gt;, the judge who gave everything a 4, has zero spread: a division by zero.&lt;/p&gt;

&lt;p&gt;The fix is shrinkage. Each judge starts with kappa imaginary projects of population-typical spread, and their own marks pull them away from it as they pile up:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;sigma_j^2 = ((n_j - 1) * s_j^2 + kappa * sigma_pop^2) / (n_j - 1 + kappa)&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;With kappa at its default of 5, the two-score judges separate by their actual gap: 0.277 for a third of a point, 0.534 for two thirds, 0.757 for a full point. That last one is above 0.707, so shrinkage can push a z-score up as well as down: when a judge's own spread is wider than the population's (0.65), shrinkage pulls it toward 0.65 and the z-score rises. And every z that &lt;code&gt;jdg_07&lt;/code&gt; produces comes out exactly 0, a neutral vote instead of a crash.&lt;/p&gt;

&lt;p&gt;My first test of this used kappa 0.001 as "plain" and failed with &lt;code&gt;AssertionError: 0.6214804438775124 not greater than 0.7&lt;/code&gt;. The test was wrong, not the engine: even a thousandth of an imaginary project pulls a nearly flat judge toward the population (0.7071 at kappa 0, 0.6215 at 0.001). The test now pins 0.7071 at kappa 0.&lt;/p&gt;

&lt;h2&gt;
  
  
  I thought plain z-scores were close enough
&lt;/h2&gt;

&lt;p&gt;The reference picked kappa 5 with no evidence, so I measured: Spearman rank correlation with a hidden true quality (1.0 means the same order), 1,000 simulated events per row, seed 2026, three reviews per project. Condensed from &lt;code&gt;python src/judging/engine.py --trials 1000&lt;/code&gt;, which gives the same numbers when rerun:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                               raw    plain z  kappa 5  kappa 25  kappa 5 beats raw
reference judges, 12 projects  0.766  0.838    0.841    0.841     75% of events
reference judges, 40 projects  0.809  0.891    0.890    0.889     98%
fixture-like, 30 judges        0.802  0.812    0.830    0.832     69%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The "plain z" column is computed at kappa 0.01, so a judge with no spread does not divide by zero; kappa 0 gives the same column to three decimals.&lt;/p&gt;

&lt;p&gt;With the reference's judges, most of the win is just centring each judge on their own mean, and any kappa from 2 to 25 lands within about 0.003 of plain z. The third row is shaped like the real fixture, 30 judges with about four whole-number scores each, and there plain z barely moves: 0.802 to 0.812. Rerun at true plain z (kappa 0), it beats the raw average in only 55% of those events (56% at the kappa 0.01 the table uses), close to a coin flip, where kappa 5 wins 69%. So plain z-scores went.&lt;/p&gt;

&lt;p&gt;Kappa 5 still fails to beat the raw average in about one fixture-like event in three, so the results page counts informative reviews (a review counts only if its judge has two or more scores with some spread) and flags projects ranked on fewer than two. Three fixture judges carry no signal, 5 of the 122 counted reviews, and one project is flagged: Small Relay. Normalization moves 33 of 40 projects, median 2 places. Dry Harbour climbs from 29th to 8th: the judge who dragged it down scored nothing else, so that lone mark normalizes to neutral, and most of its other judges marked it above their own averages (&lt;code&gt;jdg_26&lt;/code&gt; gave 4.67 against a 3.70 average).&lt;/p&gt;

&lt;h2&gt;
  
  
  I thought my rankings were deterministic
&lt;/h2&gt;

&lt;p&gt;JUDGING.md once promised that ties were broken by raw average, then project id, "so the order is always deterministic". It quoted the fixture's biggest moves from my Windows machine: Glass Beacon up 10 places, Paper Anchor down 10. Before handing judging over, I rechecked every documented figure against a fresh seed, and that was the first time the proof ran inside a Linux container. It said up 8 and down 9.&lt;/p&gt;

&lt;p&gt;Three projects average exactly 3.50, and the old code handed out positions, not ranks, so the tie went to whichever float came out a hair larger:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# before: positions 1, 2, 3, even for equal averages
&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;position&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;_rank&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reviewed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;raw_mean&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;z_mean&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pk&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
    &lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;raw_rank&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;position&lt;/span&gt;

&lt;span class="c1"&gt;# after: competition ranking (1, 2, 2, 4), values compared at 9 decimal places
&lt;/span&gt;&lt;span class="n"&gt;ordered&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;value&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;9&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Raw average 3.50&lt;/th&gt;
&lt;th&gt;Raw rank, Windows&lt;/th&gt;
&lt;th&gt;Raw rank, container&lt;/th&gt;
&lt;th&gt;Raw rank, fixed&lt;/th&gt;
&lt;th&gt;Normalized rank&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Glass Beacon&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open Kiln&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Paper Anchor&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The fix was competition ranking at nine decimals, a fixed score read order and &lt;code&gt;test_equal_scores_share_a_rank_whatever_order_they_were_summed_in&lt;/code&gt;, which took the suite from 125 tests to 126. The rebuilt container and a fresh local seed then printed the proof identically.&lt;/p&gt;

&lt;p&gt;My tests could not see it because they all ran on one machine. Working the arithmetic again, I think my docs also blame the wrong thing: they say "summation order" between two databases, but the weighted sum always walks the criteria in a fixed order. The likely culprit is Python:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Python 3.11, each criterion weighted 1/3
  marks (3, 4, 4) sum to 3.666666666666666
  marks (3, 3, 5) sum to 3.6666666666666665
  Glass Beacon's raw mean: 3.4999999999999996
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Windows venv runs 3.11, so Glass Beacon sorts last of the three despite the best z-score. The container runs 3.12.14, whose &lt;code&gt;sum()&lt;/code&gt; switched to a more accurate algorithm: all three are exactly 3.5, the old key falls through to the z-score, and Glass Beacon comes first. Both orders I saw follow from that.&lt;/p&gt;

&lt;p&gt;A day later another test failed only in the Linux image, with &lt;code&gt;AssertionError: 3.0 == 3.0&lt;/code&gt;. It compared whichever of two tied projects came first, both at 3.0 before and after the change it checked, so it had never tested what it claimed.&lt;/p&gt;

&lt;h2&gt;
  
  
  I thought quadratic voting would stop a loud minority
&lt;/h2&gt;

&lt;p&gt;The community vote was going to be quadratic, the textbook answer to a loud minority: n votes on one project cost n squared credits. My first simulation disagreed. A fifth of the voters backing one weak project won 100% of events, worse than one person, one vote at 96%, because a bloc member puts 5 votes on the target while a sincere voter's favourite gets 2 or 3. So I added a ceiling of 3 votes per project per ballot and reran with 1,000 events per row, 40 projects and 300 voters. "Seen" is how many projects a voter actually looks at:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Best project wins, nobody gaming, 10 seen&lt;/th&gt;
&lt;th&gt;Same, 20 seen&lt;/th&gt;
&lt;th&gt;20% bloc wins, 10 seen&lt;/th&gt;
&lt;th&gt;10% bloc, 3 identities each, 10 seen&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One person, one vote&lt;/td&gt;
&lt;td&gt;79%&lt;/td&gt;
&lt;td&gt;90%&lt;/td&gt;
&lt;td&gt;95%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quadratic, no ceiling&lt;/td&gt;
&lt;td&gt;64%&lt;/td&gt;
&lt;td&gt;84%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quadratic, ceiling of 3&lt;/td&gt;
&lt;td&gt;48%&lt;/td&gt;
&lt;td&gt;82%&lt;/td&gt;
&lt;td&gt;57%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At 20 projects seen, the ceiling held a 20% bloc to 0% wins where the uncapped version let it win 40%. The cost is in the first column: when nobody games the vote, one person, one vote crowns the best project (the highest hidden quality) 79% of the time at 10 seen; my capped method manages 48%. It orders the whole field better (rank correlation 0.984 against 0.926), but the winner is the number people remember.&lt;/p&gt;

&lt;p&gt;The last column is the same for every method: a bloc whose members each hold three identities wins every event. Past that point the work is in who gets a ballot: three access modes (an open link, email links stored only as hashes, or signed-in accounts) and a detector that flags bursts, shared devices and aliased addresses.&lt;/p&gt;

&lt;p&gt;Ballot order mattered too. With one shared order and voters looking at about 10 projects, 94% of the top three came from the first ten listed, and the best project won 13% of the time. A stable shuffle per voter, seeded by an HMAC of who the voter is and never stored, lifted that to 51%.&lt;/p&gt;

&lt;h2&gt;
  
  
  I thought isolation meant guarding the judges' endpoints
&lt;/h2&gt;

&lt;p&gt;By the time the T3 brief repeated the isolation requirement, judge queries started from the signed-in judge, anything outside a judge's assignments answered 404 like a missing id, and a request for a peer's scores was refused with 403 before any lookup, tested with three spellings of another judge and one id that belongs to nobody.&lt;/p&gt;

&lt;p&gt;This time the question was not what a judge can reach, but where judging data ends up. A grep for history in the templates, then a list of judging log calls that pass &lt;code&gt;team=&lt;/code&gt;, found it. In T1, the team history panel read &lt;code&gt;Activity.objects.filter(event=event, team=team)&lt;/code&gt; with no filter on the kind of entry. In T2, five judging actions (assigning by hand, unassigning, submitting a score, changing one, declaring a conflict) started tagging their audit lines with the team. Neither change was wrong alone, but together they sent T2's judging lines into a panel T1 had built to show teammates who joined and who edited.&lt;/p&gt;

&lt;p&gt;Every isolation test probed judging endpoints, and so do the checker's; this leak ran from judges to the team through a generic audit log. The seed imports the fixture's 126 scores without logging anything, which is likely why walkthroughs on seeded data showed nothing odd. The fix is one helper both pages must go through:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Entries a team must never see in its own history: who judges their project,
# when a judge scored it, a judge's conflict and their reason, and anything
# about the community vote. They stay in the organizers' activity log.
&lt;/span&gt;&lt;span class="n"&gt;STAFF_ONLY_VERBS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;score.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;judge.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;judging.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;voting.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;team_history&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;team&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;entries&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Activity&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;objects&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;team&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;team&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;prefix&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;STAFF_ONLY_VERBS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;entries&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;entries&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exclude&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;verb__startswith&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;prefix&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;entries&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select_related&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;actor&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)[:&lt;/span&gt;&lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Its test plants a conflict reason in capitals, &lt;code&gt;I MENTORED THEM&lt;/code&gt;, and asserts it reaches neither team page but still reaches the organizers' log. The weakness was built in from the start: a deny list keyed on how verbs are spelled.&lt;/p&gt;

&lt;p&gt;"Hidden looks exactly like missing" failed once more the next day, in a threat-model test comparing the API's answer for a draft with its answer for an id that does not exist:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AssertionError: {'error': 'not_found', 'detail': 'No such project, or nothing you can see.'} != {'error': 'not_found', 'detail': 'Nothing here, or nothing you can see.'}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two polite 404s, one sentence apart: enough for anyone counting up through project ids to tell a hidden draft from an empty slot. The hidden case now raises the same &lt;code&gt;Http404&lt;/code&gt; as a missing one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some of my tests pass because the attack still works
&lt;/h2&gt;

&lt;p&gt;When the tests moved out of the Django project, running the suite from inside &lt;code&gt;src/&lt;/code&gt; printed &lt;code&gt;Ran 0 tests in 0.000s OK&lt;/code&gt;: a pass that touched nothing. The test runner now defaults to the &lt;code&gt;tests/&lt;/code&gt; folder, so it cannot quietly find zero tests.&lt;/p&gt;

&lt;p&gt;The threat model, one of the two bonus challenges I took on, lists 43 attacks: 28 stopped, 3 capped or flagged, 1 logged only and 11 open. The doc and the tests share attack ids, and four of the open attacks have tests with &lt;code&gt;_gap_&lt;/code&gt; in the name that pass because the attack works:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_SY5_gap_a_fresh_browser_gets_a_fresh_ballot_on_an_open_link&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;reverse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;voting:ballot_link&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;slug&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;link_token&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;browser&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# a private window, or cookies cleared
&lt;/span&gt;        &lt;span class="n"&gt;browser&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;browser&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;reverse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;voting:api_ballot&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;slug&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?link=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;link_token&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                     &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;votes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pk&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;}}),&lt;/span&gt; &lt;span class="n"&gt;content_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;assertEqual&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Ballot&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;objects&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Close that hole and the test goes red, which forces a change to the doc that admits it. A suite that records only what works cannot tell you when your list of weaknesses has gone stale.&lt;/p&gt;

&lt;p&gt;Writing the threat model turned up four more holes besides the 404 wording. A passed deadline could be moved later under scores already given (now refused once a score or ballot exists). Submitted projects were public before the deadline, handing early ideas to teams still working (now private by default). One inbox could cast several ballots (more on that below). And nothing looked for a judge boosting a friend.&lt;/p&gt;

&lt;p&gt;For that last one I flag a project when one judge's normalized score sits far from the rest of its panel. I first picked a threshold of 2.0, because only 4 of 106 honest gaps on the fixture reached it. Then the simulation ran: at 2.0 the flag caught only 25% of single-judge boosts. The sweep, over 1,000 simulated events:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;threshold   1 colluder caught   2 colluders caught   honest projects flagged
1.0         86%                 84%                  29.7%
1.25        75%                 70%                  14.5%
1.5         58%                 56%                  5.8%
1.75        40%                 40%                  2.1%
2.0         25%                 28%                  0.6%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I settled on 1.5: it catches 58% of single boosts and flags 5.8% of honest projects, two or three in every 40, a list short enough for an organizer to read. On the real fixture the flag fires far more often: on 8 of the 32 projects it can check, against the simulated 5.8%, and the 8 projects with only two reviews can never be checked. Two colluders on a three-judge panel still get a bottom-half project into the top three in 24% of events. The cheapest defence was already there: random assignment puts a given judge on a given friend's panel about one time in ten.&lt;/p&gt;

&lt;p&gt;On its first run, my HTTP attack probe reported two attacks as working. Both were bugs in the probe, and the doc still lists them. Once fixed, the probe ran 14 attacks and the portal stopped all 14.&lt;/p&gt;

&lt;h2&gt;
  
  
  I thought an email address was a voter
&lt;/h2&gt;

&lt;p&gt;Until 29 hours into the build (for the first 20 hours that email ballots existed), an email ballot belonged to the address as typed, under a unique constraint on (event, email). My own anti-abuse detector knew better: its &lt;code&gt;canonical_email&lt;/code&gt; drops &lt;code&gt;+tags&lt;/code&gt; and Gmail dots, so &lt;code&gt;ada+2@x.org&lt;/code&gt; is &lt;code&gt;ada@x.org&lt;/code&gt;. But the ballot, the per-address link limit and the own-team check all used the typed address, so one inbox could hold one ballot per spelling, merely flagged. My docs even claimed &lt;code&gt;+tags&lt;/code&gt; were dropped everywhere.&lt;/p&gt;

&lt;p&gt;The fix touched five code paths and zero migrations. &lt;code&gt;Ballot.email&lt;/code&gt; started holding a different kind of value under the same constraint name, &lt;code&gt;one_ballot_per_email&lt;/code&gt;, with a compatibility clause in place of a data migration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;voter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;Ballot&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Kind&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;EMAIL&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# Ballots from before inboxes were used keep the address as typed.
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;qs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email__in&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;inbox&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;voter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;voter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;order_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pk&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;first&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Replaying the change shows what that shortcut costs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A ballot saved before the change as &lt;code&gt;kim+1@x.org&lt;/code&gt; is missed when the inbox returns as &lt;code&gt;kim+2@x.org&lt;/code&gt;, so one inbox gets two ballots. Only databases that held email ballots before the change are exposed; a fresh seed has none.&lt;/li&gt;
&lt;li&gt;The archive importer writes addresses raw, so &lt;code&gt;lee+1@x.org&lt;/code&gt; and &lt;code&gt;lee+2@x.org&lt;/code&gt; import as two ballots.&lt;/li&gt;
&lt;li&gt;With no stored inbox, the own-team check runs a leading-wildcard &lt;code&gt;LIKE&lt;/code&gt; and filters in Python. With 50 Gmail accounts in the table, all 50 were loaded for one check.&lt;/li&gt;
&lt;li&gt;The typed spelling is lost from the ballot (&lt;code&gt;Mal.Lory@googlemail.com&lt;/code&gt; is stored as &lt;code&gt;mallory@gmail.com&lt;/code&gt;; only the email-link row keeps it).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What I would write now: keep &lt;code&gt;email_as_typed&lt;/code&gt;, add an indexed &lt;code&gt;inbox&lt;/code&gt; column computed at write time, put the unique constraint on (event, inbox), and give &lt;code&gt;EmailPass&lt;/code&gt; and &lt;code&gt;User&lt;/code&gt; the same column. The importer could not get around a rule the database enforces. The &lt;code&gt;RunPython&lt;/code&gt; backfill would fail on any inbox already holding two ballots, and that failure is useful: it forces the "which ballot wins" decision my compatibility clause skips. One cost stays: on a provider where &lt;code&gt;a+b@&lt;/code&gt; and &lt;code&gt;a@&lt;/code&gt; are different people, they share a ballot.&lt;/p&gt;

&lt;p&gt;The second redo is the audit log. &lt;code&gt;Activity&lt;/code&gt; serves three audiences (the team, organizers, platform admins), and who may read a row is decided by its verb's spelling: of 25 verbs logged with a team, 5 are hidden only by their prefix. I would add an &lt;code&gt;audience&lt;/code&gt; column, set when the line is written and defaulting to staff, so a verb nobody thought about stays hidden instead of showing up on a team page.&lt;/p&gt;

&lt;p&gt;One choice I would keep: nothing derived is stored. The whole fixture ranking is recomputed per request in about 25 to 30 ms and 6 queries on my machine. So when I noticed, about 32 hours in, that organizers had no button to publish the judges' results, the fix was three columns on the judging settings (&lt;code&gt;results_published_at&lt;/code&gt;, &lt;code&gt;results_published_by&lt;/code&gt;, &lt;code&gt;share_feedback&lt;/code&gt;) and no stored ranking to freeze: the public page runs the same calculation as the organizers' one.&lt;/p&gt;

&lt;h2&gt;
  
  
  One rubric per event, on purpose
&lt;/h2&gt;

&lt;p&gt;The reference design allowed a rubric per track. Shipshape has one per event, and I would make that cut again. A z-score compares a judge's mark with that judge's own mean and spread, which only means something on one scale. In the fixture, 7 of the 30 judges scored in two tracks, giving 50 of the 126 scores, including &lt;code&gt;jdg_24&lt;/code&gt; and &lt;code&gt;jdg_26&lt;/code&gt;, the two judges whose session cookies &lt;code&gt;.dogfood.toml&lt;/code&gt; hands the checker. Per-track rubrics would have put 40% of the scores on two scales inside one judge's z-score. Organizers keep weights that stay editable after scoring starts, which is cheap because totals are recomputed on every request.&lt;/p&gt;

&lt;p&gt;I also skipped reliability weighting of judges, because early scores would decide how much later ones count. I did not attempt the pairwise judging bonus, and I will not dress that up as a principle: I took on the threat model and API-first instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  I thought I should claim every tier I built
&lt;/h2&gt;

&lt;p&gt;The checker has seven checks, three for T1 and four for T2, and its rule for a verified tier fits in three lines (plus a loop that drops any tier above one that failed):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;verified&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;TIERS&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;any&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tier&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;checks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ok&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;checks&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tier&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There are no T3 or T4 checks, so &lt;code&gt;any&lt;/code&gt; is false for them and no portal can get them verified. The spec says overclaiming is the one thing that costs points. I built all four tiers, and for about a day and a half my report ended &lt;code&gt;claimed T1 T2 T3 T4, verified T1 T2&lt;/code&gt;; my first commit carried it. Then I split the claims: &lt;code&gt;.dogfood.toml&lt;/code&gt; claims what a program can confirm, and the README claims T3, T4 and the two bonuses next to the tests behind them. All seven checks pass, and the report now ends:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T2  judge cannot see peer scores ...... PASS
T2  participant blocked ............... PASS
T2  csv export works .................. PASS

claimed T1 T2, verified T1 T2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I held the API docs to the same rule. Checking every &lt;code&gt;/api/v1/&lt;/code&gt; answer against its OpenAPI schema during the tests found five places the docs were wrong, and showed my 307 tests reached only 73 of 134 endpoints, so a run now fails if any endpoint goes unexercised. A Schemathesis run reached about 1,000 requests with no server error before I stopped it, and the docs say that rather than claiming a pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the failures actually were
&lt;/h2&gt;

&lt;p&gt;The formulas were the part I could prove fastest, because a simulation has no stake in my being right. The method I shipped beats the raw average in 69% of fixture-like events; plain z-scores do it in 55%. I would rather print both numbers than print a ranking and call it fair. Every serious failure here happened somewhere else: the wording of a 404, which Python summed the marks, which column a ballot was keyed on, where an audit line ended up.&lt;/p&gt;

&lt;p&gt;Checking facts for this post, I seeded a fresh fixture, set the results date three days ahead, gave "Best in show" to Glass Signal and opened the winning team's page as one of its members. Awards stay private until the results date, so organizers can decide winners before the ceremony, and the public awards list correctly returned 0 rows. The team page returned 200 with "Best in show" in its history, from this line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;log_activity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;actor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;award.given&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gave “&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;prize&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;” to “&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;”&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="n"&gt;team&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;team&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;award.&lt;/code&gt; is not on the deny list. No test covers it, and it is still there as I write this: the opening scene again, with a different verb. Adding &lt;code&gt;award.&lt;/code&gt; to the list would close this one. The &lt;code&gt;audience&lt;/code&gt; column would have kept it closed without anyone having to remember.&lt;/p&gt;

&lt;h3&gt;
  
  
  Try it
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/manusingh090/Shipshape.git
&lt;span class="nb"&gt;cd &lt;/span&gt;Shipshape
docker compose up
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No network is needed. Open &lt;a href="http://localhost:8080" rel="noopener noreferrer"&gt;http://localhost:8080&lt;/a&gt;, choose Sign in and press one of the Be buttons: Be Rosa is the organizer, Be Priya a team captain, Be Tom a participant with no team, Be Ada and Be Diego are judges. The tables come from &lt;code&gt;python src/judging/engine.py --trials 1000&lt;/code&gt; (add &lt;code&gt;--collusion&lt;/code&gt; for the collusion simulation) and &lt;code&gt;python src/voting/method.py --trials 1000&lt;/code&gt; (add &lt;code&gt;--seen 20&lt;/code&gt; for the 20-seen column and &lt;code&gt;--order&lt;/code&gt; for ballot order), which need only plain Python.&lt;/p&gt;

&lt;h3&gt;
  
  
  References
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;The DOGFOOD 2026 brief and acceptance checker, in the repo as &lt;code&gt;tests/acceptance/spec.md&lt;/code&gt; and &lt;code&gt;tests/acceptance/run.py&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;JUDGING.md in the repo: section 4 (normalization and the Monte Carlo proof), section 10 (the vote simulations), section 11 (the threat model)&lt;/li&gt;
&lt;li&gt;Schemathesis: &lt;a href="https://schemathesis.readthedocs.io/" rel="noopener noreferrer"&gt;https://schemathesis.readthedocs.io/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Shipshape: &lt;a href="https://github.com/manusingh090/Shipshape" rel="noopener noreferrer"&gt;https://github.com/manusingh090/Shipshape&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Google Slides : &lt;a href="https://docs.google.com/presentation/d/1z8AnsudKyrrzO9azSSoI51wdYGTVh7hNFrpsKxTI_Eg/edit?usp=sharing" rel="noopener noreferrer"&gt;https://docs.google.com/presentation/d/1z8AnsudKyrrzO9azSSoI51wdYGTVh7hNFrpsKxTI_Eg/edit?usp=sharing&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>django</category>
      <category>python</category>
      <category>showdev</category>
      <category>hackathonraptors</category>
    </item>
  </channel>
</rss>
