TL;DR. I built Rubrica, a self-hosted hackathon portal, for DOGFOOD 2026. The interesting part is not the feature list. It is a judging engine whose maths you can verify by hand, one result I am uncomfortable with but kept, a parameter (
k) that behaves the opposite of what you would guess, and a list of things I refused to build, with reasons.Repo: https://github.com/aanushiya170/rubrica · Python 3.12 · FastAPI · SQLite · one container · MIT
At a glance
| Fixture | 41 projects, 30 judges, 8 tracks, 40 teams, 126 scores |
| Acceptance checker | 7 checks, all T1/T2 |
| My extra checker | 31 live HTTP checks (T3 12/12, T4 19/19) |
| Test suite | 37 tests, including the layering rule and the normalization proof |
| Normalization |
k = 3 shrinkage heuristic, honestly named |
| Hash chain | Deliberately not built (see section 7) |
1. The problem that decides everything
A strict judge looks exactly like a weak project.
Suppose judge A averages 3.3 and judge B averages 3.7. Is B generous, or did B just draw better projects? If every judge sees a private batch, you cannot tell. Harshness and quality are mathematically confounded, and no formula applied afterwards can separate them.
So before touching any maths, I changed how work is assigned. Rubrica uses blind overlap:
- A judge only receives projects in tracks they cover, and never their own team's project.
- Per track,
anchor_countprojects are anchors: every eligible judge reviews them independently. - Remaining projects go to the least-loaded eligible judges until each has
reviews_per_projectreviewers. - The engine is idempotent (existing judge-project pairs are kept) and seeded (
seed=42), so regenerating gives the same plan.
The overlap is in what gets reviewed, never in who sees whose numbers. That second half is where most portals quietly fail.
Lesson 1: design the assignment before you design the maths. Normalization without overlap is just a guess with decimals.
2. The 403 that is not a hidden column
The cheap way to hide peer scores is to not render them. That is a UI decision, and the API behind it still answers.
In Rubrica, identity comes from the session and nothing else. A query parameter like ?judge= is a request that a scope check may refuse:
request -> resolve session -> route -> service function
-> scope check inside the service (assert_judge_scope / require_staff)
-> SQL -> audit.record -> signals.emit -> response
curl -i -H 'Cookie: session=<judge_b>' \
'localhost:8080/api/judge/scores?judge=jdg_24'
# HTTP/1.1 403
Because the check lives in the service layer, the HTML pages, /api and /api/v1 all return the same 403. A new route cannot forget it, because the data function itself refuses. The same applies to scoring: assert_judge_scope(project_id) checks the assignments table before any score write.
Errors are mapped to status codes in exactly one place: Unauthorized 401, Forbidden/Closed 403, NotFound 404, Conflict 409, RateLimited 429. One place means one place to audit.
Lesson 2: enforce access where the data is read, not where the page is drawn.
3. The normalization, and the name I would not use
Raw score is a weighted sum. Weights must sum to exactly 1.0 (the server rejects anything else), and every stored score carries the rubric version it was produced under, so editing a rubric never rewrites history.
S_jp = sum_c w_c * score_c(j, p) # fixture: 40% functionality, 35% quality, 25% innovation
For judge j with n_j scores, I shrink their mean and variance toward the global values with prior strength k (default 3):
mu_j* = (n_j * mu_j + k * mu_g) / (n_j + k)
sigma_j* = sqrt( (n_j * v_j + k * v_g) / (n_j + k) )
z = (S_jp - mu_j*) / sigma_j*
N_jp = clip( mu_g + z * sigma_g, 1, 5 )
A project's result is the mean of its N_jp, ties broken by project id, and the raw ranking is always displayed beside it.
This resembles empirical Bayes. It is not. It is a linear, fixed-k shrinkage heuristic, and the docs call it a sample-size-shrunk location-scale normalization heuristic, inspired by empirical-Bayes reasoning. Naming your method honestly costs one awkward sentence. Overclaiming it costs your credibility on every other number.
The worked example you can check with a calculator
Fixture globals: mean 3.568254, population sd 0.670661, variance 0.449786. Judge jdg_01 has exactly one review, a 2.00.
mu_j* = (1*2.000 + 3*3.568254) / 4 = 3.1762
sigma_j* = sqrt((1*0 + 3*0.449786) / 4) = 0.5808
z = (2.000 - 3.1762) / 0.5808 = -2.025
N = 3.568254 + (-2.025)(0.670661) = 2.210
And a judge with real data: jdg_24 reviewed 11 projects, raw mean 3.318, shrunk mean 3.372. When there is enough data, the data wins.
Why divide by n (population variance)? At n_j = 1 the sample variance is undefined and at n_j = 2 it is wildly unstable. The k * v_g term already supplies prior mass, and /n keeps sigma_j* finite and monotone in n_j.
4. The part where the maths fought back
Working through the formulas for this write-up, I found three things the docs do not say out loud.
4a. For a one-review judge, the whole pipeline collapses to one constant
When n_j = 1, the judge's variance v_j is 0. Substitute and everything simplifies:
N = mu_g + sqrt(k / (k + 1)) * (S - mu_g)
With k = 3 the factor is 0.866. A judge with a single review gets a fixed 13.4% pull toward the global mean, regardless of what they scored. Checking against the worked example: 3.568254 + 0.866 x (2.00 - 3.568254) = 2.210. It matches.
4b. Increasing k makes one-review judges pass through more untouched
This is the counterintuitive one. You would expect a stronger prior to mean stronger correction. For single-review judges it is the opposite:
k |
factor sqrt(k/(k+1))
|
a 2.00 becomes | pull toward mean |
|---|---|---|---|
| 0.5 | 0.577 | 2.663 | 42.3% |
| 1 | 0.707 | 2.459 | 29.3% |
| 2 | 0.816 | 2.288 | 18.4% |
| 3 | 0.866 | 2.210 | 13.4% |
| 5 | 0.913 | 2.137 | 8.7% |
| 10 | 0.953 | 2.073 | 4.7% |
| 20 | 0.976 | 2.038 | 2.4% |
Why: a bigger k says "assume this judge behaves like the global population". If the judge is the global population, their z-score equals the raw z-score and nothing changes. A small k says "I know almost nothing about this judge", so their lone score is treated as weak evidence and damped hard.
What this means in practice: k is not a "how much do I distrust thin judges" knob. If you want to neutralise single-review judges, raising k does the reverse. That job belongs to the low-sample flag and to averaging across reviewers.
4c. The shrinkage is milder than "a harsh score can't sink a project" sounds
A lone 2.00 becomes 2.21. Still deeply negative. On a project with 5 reviews, that judge's raw drag on the mean is -0.314; normalized, it is -0.272. The correction removes about 13% of the damage, not all of it.
So the honest claim is narrower than my first draft's: normalization softens one outlier. The real protection for a project is review count, which is why the UI shows it beside every score and flags anything under 3.
4d. How much of the fixture leans on the prior?
The own-data weight is n / (n + k). At k = 3:
judge reviews n
|
weight on judge's own data |
|---|---|
| 1 | 25% |
| 2 | 40% |
| 3 | 50% |
| 6 | 67% |
| 11 | 79% |
In the official fixture, 8 of 30 judges (27%) have two reviews or fewer, and judges with three or fewer account for 35 of the 126 reviews (28%). More than a quarter of all review mass is rated by judges for whom the prior carries half the weight or more. That is not a flaw in the method; it is the dataset, and it is why I refuse to present the normalized score without its sample size.
5. The result I am least comfortable with, and kept
The biggest fall in the fixture is prj_19: raw rank 17, normalized rank 30. It has two reviews: a 3.15 from a normal judge and a 4.00 from jdg_07.
jdg_07 gave 4/4/4 to all three of their projects. A judge who scores everything identically gives no relative separation between projects, so their 4.0 says nothing about this project versus the others. Two policies are implemented and stored on every run as zero_variance_policy:
-
baseline(default): each of their scores maps to the global mean. Contribution becomes "no information" instead of "everything is excellent". -
shrink: use the formula. All three map to the same value, 3.874. No fake ordering is invented either way.
The swing between the two is 0.305 points per review. On a two-review project that moves the mean by about 0.15. Under baseline, prj_19's only high score stops being high: raw 3.575 (the mean of 4.00 and 3.15), normalized 3.322.
Is that fair? Arguable, and I will not pretend otherwise. Here are the largest movers:
| project | reviews | raw avg | normalized | raw rank | normalized rank |
|---|---|---|---|---|---|
| prj_19 | 2 | 3.575 | 3.322 | 17 | 30 |
| prj_28 | 3 | 3.400 | 3.108 | 28 | 37 |
| prj_27 | 3 | 3.383 | 3.507 | 29 | 21 |
| prj_12 | 3 | 3.450 | 3.553 | 23 | 16 |
The top of the table is stable: prj_34 and prj_11 are first and second both ways. That is what you want from calibration. It should move the borderline, not overturn the obvious.
Lesson 3: make judgement calls into stored, visible parameters. k and zero_variance_policy are recorded on each run, so the argument about fairness can be had with data instead of vibes.
6. Pairwise judging: Bradley-Terry in about forty lines
As a bonus, a judge can pick the stronger of two assigned projects (extensions/pairwise.py). Strengths are fitted with the MM algorithm (Hunter 2004):
P(i beats j) = pi_i / (pi_i + pi_j)
pi_i <- W_i / sum_{j != i} n_ij / (pi_i + pi_j)
Two details matter. Every pair gets 0.5 symmetric pseudo-wins, so a project with no comparisons still has a finite estimate under sparse data. Strengths are normalised to geometric mean 1. Wins and comparison counts are shown beside each strength, for the same reason review counts are shown beside normalized scores.
I deliberately did not blend pairwise into the rubric ranking. Two ranking methods with different assumptions averaged together produce a number nobody can explain.
7. Publishing freezes inputs, and why there is no hash chain
When an organizer publishes, result_snapshots stores the exact score_ids, project_ids, exclusions and parameters, plus a SHA-256 of those inputs. Verify recomputes the whole ranking from the recorded evidence and diffs every published row. A test tampers with a score and asserts verification fails.
I did not build a hash-chained audit log, and the README says so. A chain stored in the same database as the thing it protects has the same trust boundary: anyone who can rewrite the rows can rewrite the chain. The claim I can honestly make and test is recomputation from recorded inputs, plus the advice to export audit.csv and the snapshot JSON off-box after publishing.
The audit log is still append-only and monotonic by seq, and every record is also emitted as a signal. That signal bus is the only bridge between the layers, which is the next decision.
Lesson 4: build the claim you can test, not the one that sounds strongest.
8. One rule that kept four tiers from tangling
Everything sits on a T1+T2 core that must never break.
src/rubrica/core/never importssrc/rubrica/extensions/.
tests/test_layering.py enforces it, and also boots the app with RUBRICA_EXTENSIONS=0 to prove the core runs alone. Delete extensions/ and the seven acceptance checks still pass.
The T4 extras (scoped-key REST API with an OpenAPI 3.1 document, HMAC-signed webhooks, Ed25519-signed judge records verifiable offline, embeddable gallery, bulk import/export, pairwise judging) attach through register(app) and optional startup hooks. Core never calls them. Webhooks listen: they subscribe to the audit signal stream with *.
Small calls that paid off:
-
Two seeded events.
evt_01is the official fixture, loaded verbatim and markedis_historical(read-only), never mutated by the seed.evt_02is a live demo event with open submissions, so the submit-judge-publish lifecycle can be exercised on a deadline that is not already in the past. - Idempotent seed. Booting twice still yields exactly 41 fixture projects, asserted in the boot log and in tests.
-
Per-voter ballot order is a deterministic shuffle seeded by
sha256(event:voter): random across voters, stable across reloads for one voter, so position bias does not favour the same project for everyone. -
Duplicates are resolved, never deleted.
canonical_project_idplus statusexcluded; the scores stay as evidence and ranking drops the excluded record. -
Votes use a
UNIQUEconstraint, not only application code, and the rejected duplicate is itself audited asvote.duplicate_rejected.
9. When the spec was harder than it looked: the checker checks half my claim
run.py has seven checks and all seven are T1/T2. If a portal claims T3 or T4, it prints "claimed but not verified" no matter what. That is correct for a checker, but it makes a T3/T4 claim unfalsifiable.
So I wrote check_extended.py: 31 more live HTTP checks in the same style (stdlib only, one PASS/FAIL line each), with output committed as acceptance-report-extended.txt:
T3: 12/12 verified
T4: 19/19 verified
Lesson 5: if you claim a feature, ship a command that can prove you wrong.
10. Features I cut, and do not regret
| Cut | Why |
|---|---|
| Hash-chained audit log | Same trust boundary as the data (section 7) |
| Outbound email | Offline rule; invite links are shown to the organizer instead |
| Blended pairwise + rubric ranking | Two methods averaged into one unexplainable number |
| Fancy front end | Server-rendered Jinja, inline CSS, no build step, no CDN; complete for every flow, deliberately not the differentiator |
| Auth-as-a-service | Hand-rolled scrypt passwords and opaque HttpOnly session tokens; the checker needs only a header |
11. What I would redo
Listed in the repo before anyone asked, because a gap that is a decision beats a gap that is a surprise:
- The rate limiter is in-process. Right for one container, wrong for two. Next: move it into the database or Redis.
- No email verification. Sybil resistance rests on accounts, a 5-vote quota, per-IP limits and an audit trail. A determined attacker can still create accounts, and the threat model says so.
-
No per-form CSRF token.
SameSite=Laxplus non-GET mutations blocks cross-site form posts, but a same-site XSS would not be contained. - Webhooks retry zero times. Failures are recorded with the error; there is a manual test endpoint but no backoff queue.
- Location-scale only. I correct a judge's average level and spread, not judge-by-track interactions, criterion-specific harshness or drift over time.
- No real migration runner. The table exists; only the base version is used.
- Collusion is only partly covered. Blind overlap and shrinkage help, but coordinated, consistent inflation across several judges is not detectable from scores alone. The calibration table and per-judge raw means are the tool for that conversation.
Five takeaways
- Overlap makes calibration possible. Without anchors, judge harshness and project quality are confounded.
- Enforce access in the service, not the template. Then every surface inherits it.
- Do not oversell your statistics. "Shrinkage heuristic, k=3, here are its limits" survives scrutiny. "Bayesian" does not.
-
Check your parameters' behaviour, not their names.
ksounds like "distrust", and for thin judges it does the opposite. - Claim only what you can falsify. Recomputation beats a ceremonial chain, and a second checker beats an unverifiable tier.
Try it
git clone https://github.com/aanushiya170/rubrica
cd rubrica
docker compose up
python3 run.py .dogfood.toml
python3 check_extended.py .dogfood.toml
Built for DOGFOOD 2026 with @hackathonraptors. If you disagree with a normalization choice, especially the baseline zero-variance policy, I would genuinely like to hear it.
Top comments (0)