<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sasi Sanjay</title>
    <description>The latest articles on DEV Community by Sasi Sanjay (@sasi_sanjay_9fff84a9df184).</description>
    <link>https://dev.to/sasi_sanjay_9fff84a9df184</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4117614%2F1f027e24-6c71-4653-8fc1-8b1f1c4109de.jpg</url>
      <title>DEV Community: Sasi Sanjay</title>
      <link>https://dev.to/sasi_sanjay_9fff84a9df184</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sasi_sanjay_9fff84a9df184"/>
    <language>en</language>
    <item>
      <title>Normalizing 30 Judges Without Lying About It: the maths, a counterintuitive k, and what I chose not to build</title>
      <dc:creator>Sasi Sanjay</dc:creator>
      <pubDate>Fri, 02 Oct 2026 10:17:59 +0000</pubDate>
      <link>https://dev.to/sasi_sanjay_9fff84a9df184/normalizing-30-judges-without-lying-about-it-the-maths-a-counterintuitive-k-and-what-i-chose-not-nb1</link>
      <guid>https://dev.to/sasi_sanjay_9fff84a9df184/normalizing-30-judges-without-lying-about-it-the-maths-a-counterintuitive-k-and-what-i-chose-not-nb1</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR.&lt;/strong&gt; I built &lt;strong&gt;Rubrica&lt;/strong&gt;, a self-hosted hackathon portal, for DOGFOOD 2026. The interesting part is not the feature list. It is a judging engine whose maths you can verify by hand, one result I am uncomfortable with but kept, a parameter (&lt;code&gt;k&lt;/code&gt;) that behaves the opposite of what you would guess, and a list of things I refused to build, with reasons.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/aanushiya170/rubrica" rel="noopener noreferrer"&gt;https://github.com/aanushiya170/rubrica&lt;/a&gt; · Python 3.12 · FastAPI · SQLite · one container · MIT&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  At a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fixture&lt;/td&gt;
&lt;td&gt;41 projects, 30 judges, 8 tracks, 40 teams, 126 scores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Acceptance checker&lt;/td&gt;
&lt;td&gt;7 checks, all T1/T2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;My extra checker&lt;/td&gt;
&lt;td&gt;31 live HTTP checks (T3 12/12, T4 19/19)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Test suite&lt;/td&gt;
&lt;td&gt;37 tests, including the layering rule and the normalization proof&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Normalization&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;k = 3&lt;/code&gt; shrinkage heuristic, honestly named&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hash chain&lt;/td&gt;
&lt;td&gt;Deliberately not built (see section 7)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  1. The problem that decides everything
&lt;/h2&gt;

&lt;p&gt;A strict judge looks &lt;strong&gt;exactly&lt;/strong&gt; like a weak project.&lt;/p&gt;

&lt;p&gt;Suppose judge A averages 3.3 and judge B averages 3.7. Is B generous, or did B just draw better projects? If every judge sees a private batch, you cannot tell. Harshness and quality are mathematically confounded, and no formula applied afterwards can separate them.&lt;/p&gt;

&lt;p&gt;So before touching any maths, I changed how work is &lt;strong&gt;assigned&lt;/strong&gt;. Rubrica uses &lt;strong&gt;blind overlap&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A judge only receives projects in tracks they cover, and never their own team's project.&lt;/li&gt;
&lt;li&gt;Per track, &lt;code&gt;anchor_count&lt;/code&gt; projects are &lt;strong&gt;anchors&lt;/strong&gt;: every eligible judge reviews them independently.&lt;/li&gt;
&lt;li&gt;Remaining projects go to the least-loaded eligible judges until each has &lt;code&gt;reviews_per_project&lt;/code&gt; reviewers.&lt;/li&gt;
&lt;li&gt;The engine is idempotent (existing judge-project pairs are kept) and seeded (&lt;code&gt;seed=42&lt;/code&gt;), so regenerating gives the same plan.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The overlap is in &lt;strong&gt;what&lt;/strong&gt; gets reviewed, never in &lt;strong&gt;who sees whose numbers&lt;/strong&gt;. That second half is where most portals quietly fail.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lesson 1:&lt;/strong&gt; design the assignment before you design the maths. Normalization without overlap is just a guess with decimals.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The 403 that is not a hidden column
&lt;/h2&gt;

&lt;p&gt;The cheap way to hide peer scores is to not render them. That is a UI decision, and the API behind it still answers.&lt;/p&gt;

&lt;p&gt;In Rubrica, identity comes from the session and nothing else. A query parameter like &lt;code&gt;?judge=&lt;/code&gt; is a &lt;em&gt;request&lt;/em&gt; that a scope check may refuse:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;request -&amp;gt; resolve session -&amp;gt; route -&amp;gt; service function
        -&amp;gt; scope check inside the service (assert_judge_scope / require_staff)
        -&amp;gt; SQL -&amp;gt; audit.record -&amp;gt; signals.emit -&amp;gt; response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Cookie: session=&amp;lt;judge_b&amp;gt;'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'localhost:8080/api/judge/scores?judge=jdg_24'&lt;/span&gt;
&lt;span class="c"&gt;# HTTP/1.1 403&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because the check lives in the &lt;strong&gt;service layer&lt;/strong&gt;, the HTML pages, &lt;code&gt;/api&lt;/code&gt; and &lt;code&gt;/api/v1&lt;/code&gt; all return the same 403. A new route cannot forget it, because the data function itself refuses. The same applies to scoring: &lt;code&gt;assert_judge_scope(project_id)&lt;/code&gt; checks the assignments table before any score write.&lt;/p&gt;

&lt;p&gt;Errors are mapped to status codes in exactly one place: Unauthorized 401, Forbidden/Closed 403, NotFound 404, Conflict 409, RateLimited 429. One place means one place to audit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lesson 2:&lt;/strong&gt; enforce access where the data is read, not where the page is drawn.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The normalization, and the name I would not use
&lt;/h2&gt;

&lt;p&gt;Raw score is a weighted sum. Weights must sum to exactly 1.0 (the server rejects anything else), and every stored score carries the rubric version it was produced under, so editing a rubric never rewrites history.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;S_jp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sum_c&lt;/span&gt;  &lt;span class="n"&gt;w_c&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nf"&gt;score_c&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="c1"&gt;# fixture: 40% functionality, 35% quality, 25% innovation
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For judge &lt;code&gt;j&lt;/code&gt; with &lt;code&gt;n_j&lt;/code&gt; scores, I shrink their mean and variance toward the global values with prior strength &lt;code&gt;k&lt;/code&gt; (default 3):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mu_j*    = (n_j * mu_j + k * mu_g) / (n_j + k)
sigma_j* = sqrt( (n_j * v_j + k * v_g) / (n_j + k) )

z    = (S_jp - mu_j*) / sigma_j*
N_jp = clip( mu_g + z * sigma_g, 1, 5 )
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A project's result is the mean of its &lt;code&gt;N_jp&lt;/code&gt;, ties broken by project id, and the raw ranking is always displayed beside it.&lt;/p&gt;

&lt;p&gt;This resembles empirical Bayes. &lt;strong&gt;It is not.&lt;/strong&gt; It is a linear, fixed-&lt;code&gt;k&lt;/code&gt; shrinkage heuristic, and the docs call it a &lt;em&gt;sample-size-shrunk location-scale normalization heuristic, inspired by empirical-Bayes reasoning&lt;/em&gt;. Naming your method honestly costs one awkward sentence. Overclaiming it costs your credibility on every other number.&lt;/p&gt;

&lt;h3&gt;
  
  
  The worked example you can check with a calculator
&lt;/h3&gt;

&lt;p&gt;Fixture globals: mean &lt;strong&gt;3.568254&lt;/strong&gt;, population sd &lt;strong&gt;0.670661&lt;/strong&gt;, variance &lt;strong&gt;0.449786&lt;/strong&gt;. Judge &lt;code&gt;jdg_01&lt;/code&gt; has exactly one review, a 2.00.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mu_j*    = (1*2.000 + 3*3.568254) / 4          = 3.1762
sigma_j* = sqrt((1*0 + 3*0.449786) / 4)        = 0.5808
z        = (2.000 - 3.1762) / 0.5808           = -2.025
N        = 3.568254 + (-2.025)(0.670661)       = 2.210
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And a judge with real data: &lt;code&gt;jdg_24&lt;/code&gt; reviewed 11 projects, raw mean 3.318, shrunk mean 3.372. When there is enough data, the data wins.&lt;/p&gt;

&lt;p&gt;Why divide by &lt;code&gt;n&lt;/code&gt; (population variance)? At &lt;code&gt;n_j = 1&lt;/code&gt; the sample variance is undefined and at &lt;code&gt;n_j = 2&lt;/code&gt; it is wildly unstable. The &lt;code&gt;k * v_g&lt;/code&gt; term already supplies prior mass, and &lt;code&gt;/n&lt;/code&gt; keeps &lt;code&gt;sigma_j*&lt;/code&gt; finite and monotone in &lt;code&gt;n_j&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. The part where the maths fought back
&lt;/h2&gt;

&lt;p&gt;Working through the formulas for this write-up, I found three things the docs do not say out loud.&lt;/p&gt;

&lt;h3&gt;
  
  
  4a. For a one-review judge, the whole pipeline collapses to one constant
&lt;/h3&gt;

&lt;p&gt;When &lt;code&gt;n_j = 1&lt;/code&gt;, the judge's variance &lt;code&gt;v_j&lt;/code&gt; is 0. Substitute and everything simplifies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;N = mu_g + sqrt(k / (k + 1)) * (S - mu_g)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;code&gt;k = 3&lt;/code&gt; the factor is &lt;strong&gt;0.866&lt;/strong&gt;. A judge with a single review gets a fixed 13.4% pull toward the global mean, &lt;em&gt;regardless of what they scored&lt;/em&gt;. Checking against the worked example: 3.568254 + 0.866 x (2.00 - 3.568254) = 2.210. It matches.&lt;/p&gt;

&lt;h3&gt;
  
  
  4b. Increasing &lt;code&gt;k&lt;/code&gt; makes one-review judges pass through &lt;em&gt;more&lt;/em&gt; untouched
&lt;/h3&gt;

&lt;p&gt;This is the counterintuitive one. You would expect a stronger prior to mean stronger correction. For single-review judges it is the opposite:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;code&gt;k&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;factor &lt;code&gt;sqrt(k/(k+1))&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;a 2.00 becomes&lt;/th&gt;
&lt;th&gt;pull toward mean&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0.5&lt;/td&gt;
&lt;td&gt;0.577&lt;/td&gt;
&lt;td&gt;2.663&lt;/td&gt;
&lt;td&gt;42.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0.707&lt;/td&gt;
&lt;td&gt;2.459&lt;/td&gt;
&lt;td&gt;29.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0.816&lt;/td&gt;
&lt;td&gt;2.288&lt;/td&gt;
&lt;td&gt;18.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.866&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.210&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;13.4%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;0.913&lt;/td&gt;
&lt;td&gt;2.137&lt;/td&gt;
&lt;td&gt;8.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;0.953&lt;/td&gt;
&lt;td&gt;2.073&lt;/td&gt;
&lt;td&gt;4.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;0.976&lt;/td&gt;
&lt;td&gt;2.038&lt;/td&gt;
&lt;td&gt;2.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Why: a bigger &lt;code&gt;k&lt;/code&gt; says "assume this judge behaves like the global population". If the judge &lt;em&gt;is&lt;/em&gt; the global population, their z-score equals the raw z-score and nothing changes. A small &lt;code&gt;k&lt;/code&gt; says "I know almost nothing about this judge", so their lone score is treated as weak evidence and damped hard.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What this means in practice:&lt;/strong&gt; &lt;code&gt;k&lt;/code&gt; is not a "how much do I distrust thin judges" knob. If you want to neutralise single-review judges, raising &lt;code&gt;k&lt;/code&gt; does the reverse. That job belongs to the &lt;strong&gt;low-sample flag&lt;/strong&gt; and to averaging across reviewers.&lt;/p&gt;

&lt;h3&gt;
  
  
  4c. The shrinkage is milder than "a harsh score can't sink a project" sounds
&lt;/h3&gt;

&lt;p&gt;A lone 2.00 becomes 2.21. Still deeply negative. On a project with 5 reviews, that judge's raw drag on the mean is -0.314; normalized, it is -0.272. The correction removes about 13% of the damage, not all of it.&lt;/p&gt;

&lt;p&gt;So the honest claim is narrower than my first draft's: normalization &lt;em&gt;softens&lt;/em&gt; one outlier. The real protection for a project is &lt;strong&gt;review count&lt;/strong&gt;, which is why the UI shows it beside every score and flags anything under 3.&lt;/p&gt;

&lt;h3&gt;
  
  
  4d. How much of the fixture leans on the prior?
&lt;/h3&gt;

&lt;p&gt;The own-data weight is &lt;code&gt;n / (n + k)&lt;/code&gt;. At &lt;code&gt;k = 3&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;judge reviews &lt;code&gt;n&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;weight on judge's own data&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;40%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;50%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;67%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;79%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In the official fixture, &lt;strong&gt;8 of 30 judges (27%) have two reviews or fewer&lt;/strong&gt;, and judges with three or fewer account for &lt;strong&gt;35 of the 126 reviews (28%)&lt;/strong&gt;. More than a quarter of all review mass is rated by judges for whom the prior carries half the weight or more. That is not a flaw in the method; it is the dataset, and it is why I refuse to present the normalized score without its sample size.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The result I am least comfortable with, and kept
&lt;/h2&gt;

&lt;p&gt;The biggest fall in the fixture is &lt;code&gt;prj_19&lt;/code&gt;: raw rank 17, normalized rank &lt;strong&gt;30&lt;/strong&gt;. It has two reviews: a 3.15 from a normal judge and a 4.00 from &lt;code&gt;jdg_07&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;jdg_07&lt;/code&gt; gave &lt;strong&gt;4/4/4 to all three&lt;/strong&gt; of their projects. A judge who scores everything identically gives no relative separation between projects, so their 4.0 says nothing about &lt;em&gt;this&lt;/em&gt; project versus the others. Two policies are implemented and stored on every run as &lt;code&gt;zero_variance_policy&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;baseline&lt;/code&gt; (default):&lt;/strong&gt; each of their scores maps to the global mean. Contribution becomes "no information" instead of "everything is excellent".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;shrink&lt;/code&gt;:&lt;/strong&gt; use the formula. All three map to the same value, &lt;strong&gt;3.874&lt;/strong&gt;. No fake ordering is invented either way.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The swing between the two is 0.305 points per review. On a two-review project that moves the mean by about &lt;strong&gt;0.15&lt;/strong&gt;. Under &lt;code&gt;baseline&lt;/code&gt;, &lt;code&gt;prj_19&lt;/code&gt;'s only high score stops being high: raw 3.575 (the mean of 4.00 and 3.15), normalized 3.322.&lt;/p&gt;

&lt;p&gt;Is that fair? Arguable, and I will not pretend otherwise. Here are the largest movers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;project&lt;/th&gt;
&lt;th&gt;reviews&lt;/th&gt;
&lt;th&gt;raw avg&lt;/th&gt;
&lt;th&gt;normalized&lt;/th&gt;
&lt;th&gt;raw rank&lt;/th&gt;
&lt;th&gt;normalized rank&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;prj_19&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;3.575&lt;/td&gt;
&lt;td&gt;3.322&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prj_28&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;3.400&lt;/td&gt;
&lt;td&gt;3.108&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prj_27&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;3.383&lt;/td&gt;
&lt;td&gt;3.507&lt;/td&gt;
&lt;td&gt;29&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prj_12&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;3.450&lt;/td&gt;
&lt;td&gt;3.553&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The top of the table is stable: &lt;code&gt;prj_34&lt;/code&gt; and &lt;code&gt;prj_11&lt;/code&gt; are first and second &lt;strong&gt;both ways&lt;/strong&gt;. That is what you want from calibration. It should move the borderline, not overturn the obvious.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lesson 3:&lt;/strong&gt; make judgement calls into stored, visible parameters. &lt;code&gt;k&lt;/code&gt; and &lt;code&gt;zero_variance_policy&lt;/code&gt; are recorded on each run, so the argument about fairness can be had with data instead of vibes.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Pairwise judging: Bradley-Terry in about forty lines
&lt;/h2&gt;

&lt;p&gt;As a bonus, a judge can pick the stronger of two assigned projects (&lt;code&gt;extensions/pairwise.py&lt;/code&gt;). Strengths are fitted with the MM algorithm (Hunter 2004):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight r"&gt;&lt;code&gt;&lt;span class="n"&gt;P&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;beats&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;pi_i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pi_i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;pi_j&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;pi_i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;lt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;W_i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;sum_&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;!=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;n_ij&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pi_i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;pi_j&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details matter. Every pair gets &lt;strong&gt;0.5 symmetric pseudo-wins&lt;/strong&gt;, so a project with no comparisons still has a finite estimate under sparse data. Strengths are normalised to geometric mean 1. Wins and comparison counts are shown beside each strength, for the same reason review counts are shown beside normalized scores.&lt;/p&gt;

&lt;p&gt;I deliberately did &lt;strong&gt;not&lt;/strong&gt; blend pairwise into the rubric ranking. Two ranking methods with different assumptions averaged together produce a number nobody can explain.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Publishing freezes inputs, and why there is no hash chain
&lt;/h2&gt;

&lt;p&gt;When an organizer publishes, &lt;code&gt;result_snapshots&lt;/code&gt; stores the exact &lt;code&gt;score_ids&lt;/code&gt;, &lt;code&gt;project_ids&lt;/code&gt;, exclusions and parameters, plus a &lt;strong&gt;SHA-256 of those inputs&lt;/strong&gt;. &lt;strong&gt;Verify&lt;/strong&gt; recomputes the whole ranking from the recorded evidence and diffs every published row. A test tampers with a score and asserts verification fails.&lt;/p&gt;

&lt;p&gt;I did &lt;strong&gt;not&lt;/strong&gt; build a hash-chained audit log, and the README says so. A chain stored in the same database as the thing it protects has the &lt;strong&gt;same trust boundary&lt;/strong&gt;: anyone who can rewrite the rows can rewrite the chain. The claim I can honestly make and test is &lt;em&gt;recomputation from recorded inputs&lt;/em&gt;, plus the advice to export &lt;code&gt;audit.csv&lt;/code&gt; and the snapshot JSON off-box after publishing.&lt;/p&gt;

&lt;p&gt;The audit log is still append-only and monotonic by &lt;code&gt;seq&lt;/code&gt;, and every record is also emitted as a signal. That signal bus is the only bridge between the layers, which is the next decision.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lesson 4:&lt;/strong&gt; build the claim you can test, not the one that sounds strongest.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. One rule that kept four tiers from tangling
&lt;/h2&gt;

&lt;p&gt;Everything sits on a T1+T2 core that must never break.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;src/rubrica/core/&lt;/code&gt; never imports &lt;code&gt;src/rubrica/extensions/&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;code&gt;tests/test_layering.py&lt;/code&gt; enforces it, and also boots the app with &lt;code&gt;RUBRICA_EXTENSIONS=0&lt;/code&gt; to prove the core runs alone. Delete &lt;code&gt;extensions/&lt;/code&gt; and the seven acceptance checks still pass.&lt;/p&gt;

&lt;p&gt;The T4 extras (scoped-key REST API with an OpenAPI 3.1 document, HMAC-signed webhooks, Ed25519-signed judge records verifiable offline, embeddable gallery, bulk import/export, pairwise judging) attach through &lt;code&gt;register(app)&lt;/code&gt; and optional startup hooks. Core never calls them. Webhooks &lt;strong&gt;listen&lt;/strong&gt;: they subscribe to the audit signal stream with &lt;code&gt;*&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Small calls that paid off:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Two seeded events.&lt;/strong&gt; &lt;code&gt;evt_01&lt;/code&gt; is the official fixture, loaded verbatim and marked &lt;code&gt;is_historical&lt;/code&gt; (read-only), never mutated by the seed. &lt;code&gt;evt_02&lt;/code&gt; is a live demo event with open submissions, so the submit-judge-publish lifecycle can be exercised on a deadline that is not already in the past.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Idempotent seed.&lt;/strong&gt; Booting twice still yields exactly 41 fixture projects, asserted in the boot log and in tests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-voter ballot order&lt;/strong&gt; is a deterministic shuffle seeded by &lt;code&gt;sha256(event:voter)&lt;/code&gt;: random across voters, stable across reloads for one voter, so position bias does not favour the same project for everyone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Duplicates are resolved, never deleted.&lt;/strong&gt; &lt;code&gt;canonical_project_id&lt;/code&gt; plus status &lt;code&gt;excluded&lt;/code&gt;; the scores stay as evidence and ranking drops the excluded record.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Votes use a &lt;code&gt;UNIQUE&lt;/code&gt; constraint&lt;/strong&gt;, not only application code, and the rejected duplicate is itself audited as &lt;code&gt;vote.duplicate_rejected&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  9. When the spec was harder than it looked: the checker checks half my claim
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;run.py&lt;/code&gt; has seven checks and all seven are T1/T2. If a portal claims T3 or T4, it prints "claimed but not verified" no matter what. That is correct for a checker, but it makes a T3/T4 claim unfalsifiable.&lt;/p&gt;

&lt;p&gt;So I wrote &lt;code&gt;check_extended.py&lt;/code&gt;: 31 more live HTTP checks in the same style (stdlib only, one PASS/FAIL line each), with output committed as &lt;code&gt;acceptance-report-extended.txt&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T3: 12/12 verified
T4: 19/19 verified
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Lesson 5:&lt;/strong&gt; if you claim a feature, ship a command that can prove you wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Features I cut, and do not regret
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cut&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hash-chained audit log&lt;/td&gt;
&lt;td&gt;Same trust boundary as the data (section 7)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Outbound email&lt;/td&gt;
&lt;td&gt;Offline rule; invite links are shown to the organizer instead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blended pairwise + rubric ranking&lt;/td&gt;
&lt;td&gt;Two methods averaged into one unexplainable number&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fancy front end&lt;/td&gt;
&lt;td&gt;Server-rendered Jinja, inline CSS, no build step, no CDN; complete for every flow, deliberately not the differentiator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Auth-as-a-service&lt;/td&gt;
&lt;td&gt;Hand-rolled scrypt passwords and opaque HttpOnly session tokens; the checker needs only a header&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  11. What I would redo
&lt;/h2&gt;

&lt;p&gt;Listed in the repo before anyone asked, because a gap that is a decision beats a gap that is a surprise:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The rate limiter is in-process.&lt;/strong&gt; Right for one container, wrong for two. Next: move it into the database or Redis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No email verification.&lt;/strong&gt; Sybil resistance rests on accounts, a 5-vote quota, per-IP limits and an audit trail. A determined attacker can still create accounts, and the threat model says so.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No per-form CSRF token.&lt;/strong&gt; &lt;code&gt;SameSite=Lax&lt;/code&gt; plus non-GET mutations blocks cross-site form posts, but a same-site XSS would not be contained.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Webhooks retry zero times.&lt;/strong&gt; Failures are recorded with the error; there is a manual test endpoint but no backoff queue.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Location-scale only.&lt;/strong&gt; I correct a judge's average level and spread, not judge-by-track interactions, criterion-specific harshness or drift over time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No real migration runner.&lt;/strong&gt; The table exists; only the base version is used.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Collusion is only partly covered.&lt;/strong&gt; Blind overlap and shrinkage help, but coordinated, consistent inflation across several judges is not detectable from scores alone. The calibration table and per-judge raw means are the tool for that conversation.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Five takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Overlap makes calibration possible.&lt;/strong&gt; Without anchors, judge harshness and project quality are confounded.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enforce access in the service, not the template.&lt;/strong&gt; Then every surface inherits it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not oversell your statistics.&lt;/strong&gt; "Shrinkage heuristic, k=3, here are its limits" survives scrutiny. "Bayesian" does not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check your parameters' behaviour, not their names.&lt;/strong&gt; &lt;code&gt;k&lt;/code&gt; sounds like "distrust", and for thin judges it does the opposite.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claim only what you can falsify.&lt;/strong&gt; Recomputation beats a ceremonial chain, and a second checker beats an unverifiable tier.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/aanushiya170/rubrica
&lt;span class="nb"&gt;cd &lt;/span&gt;rubrica
docker compose up
python3 run.py .dogfood.toml
python3 check_extended.py .dogfood.toml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Built for DOGFOOD 2026 with @hackathonraptors. If you disagree with a normalization choice, especially the &lt;code&gt;baseline&lt;/code&gt; zero-variance policy, I would genuinely like to hear it.&lt;/p&gt;

</description>
      <category>hackathonraptors</category>
      <category>python</category>
      <category>fastapi</category>
      <category>sqlite</category>
    </item>
  </channel>
</rss>
