My portal answers all seven of the acceptance checker's questions correctly. The report is committed in the repo: seven PASS lines and a tidy last line that says claimed T1 T2, verified T1 T2.
After the freeze I did the kind of thing that only makes sense once a hackathon is over. I copied the project, put one old bug back in (the exact lines from my own commit history), and ran the checker again.
T2 judge sees own scores ............. PASS
T2 judge cannot see peer scores ...... PASS
T2 participant blocked ............... PASS
T2 csv export works .................. PASS
claimed T1 T2, verified T1 T2
Same tidy last line. Meanwhile, in that same running copy, a judge who owns exactly one ballot asked for her own scores and received two.
Nadia asks for her OWN scores and receives 2 ballot(s):
Nadia's private note: loved it
Omar's private note: meh
That is the first half of this post: what a green report is allowed to mean. The second half is about bugs of a different species, the kind you only find by running the shipped thing like a stranger would. I did not do that until it was far too late, and it is my biggest regret of the weekend.
TL;DR
- Rubric is a self-hostable hackathon submission and judging portal (Django, PostgreSQL, one
docker compose up), built for DogFood 2026, whose brief is "build the platform that will judge you". It claims tiers T1 and T2.- The acceptance checker asks seven questions. I broke a scratch copy of my own code five ways: the checker noticed 2 of 5, my own tests noticed 5 of 5.
- Bugs and their lifespans: an isolation leak (about an hour), a misspelled field that hid the live link from every judge (well over a day), a deliberately leaky debug route (around 45 hours), and a container that served no CSS or JavaScript while CI and the checker stayed green.
- The front door of my own app was undocumented, with a README pointing at a password that did not exist, because nothing I wrote ever tried it.
- Everything reproduces from one commit,
777bbb5, plus two small scripts.
Here is the thing being judged. Rubric is Django 5.2 and PostgreSQL 16 with server-rendered pages and HTMX; one docker compose up gives a seeded portal on localhost:8080 with the network off. It has five roles, events, teams by invite link, submissions with a deadline that holds, a public gallery, expiring judge invitations, an organizer-weighted rubric, seeded judge assignment, cross-judge normalization and a live organizer dashboard, with every state change written to a hash-chained audit log. The voting tier is built but cannot be machine-verified, so I claim T1 and T2 only.
This is the judge's ballot: one project, three weighted criteria, keyboard shortcuts, autosave on every click.
72 hours in one picture
My plan had a lovely shape. The original window was the 25th to the 28th in my time: both weekend days in full and exactly one weekday. Then the dates shifted by a day (the 26th to the 29th), and the same 72 hours now held two weekdays and one fewer weekend day. Classes did not get the memo.
Disrupted schedule, though: is that not the job description? Engineers build inside constraints they did not choose, and the constraint is half the fun. I had a gate-by-gate plan written before kickoff (the grey strip below), and my commit history is a rather honest record of what happened to it once the calendar changed.
Reading the chart like a detective:
- Ahead of plan early. Gate G5 was due around hours 42 to 50. It landed around hour 39.
- Then a weekday happened. From Monday afternoon to early Tuesday morning, my time, there is not a single commit, about 15 hours. College does not pause for hackathons.
- The cheap stretch items never got their slot. G7 (around hours 58 to 64) was reserved for API parity, import and export, and pairwise mode "only if genuinely ahead". It never started, because the gate before it, voting plus hardening, ran until around hour 63 and ate it.
- Tests nearly matched the application. Roughly 7,600 lines of tests against 9,500 of application code, plus 4,500 of templates and CSS, about 90% of which landed in the first 39 hours or so. Then the UI stopped changing around hour 63, with roughly nine hours still on the clock. (Lines added, not net; treat the proportions as a mood.)
So the screens existed early. What I kept postponing was looking at them. Later is a word that does not appear in a 72-hour timetable, and I will come back to what it cost.
Seven questions, and what a PASS is allowed to mean
DogFood gives every team the same small program, run.py. It never logs in (you hand it one header per role) and it makes exactly seven requests against your running portal:
| # | Tier | Request | It passes when |
|---|---|---|---|
| 1 | T1 |
GET /projects, no auth |
200 |
| 2 | T1 | same page | one of the first three fixture project titles appears in the body |
| 3 | T1 |
POST /projects/new as a participant |
any 4xx (the fixture event closed in the past) |
| 4 | T2 | judge A asks for judge scores | 200 |
| 5 | T2 | judge B asks for judge A's scores | 401 or 403 |
| 6 | T2 | a participant asks for judge scores | 401 or 403 |
| 7 | T2 | organizer asks for the CSV | 200 and a comma in the first line |
Now read that table like an adversary. Every judge in it is a fixture judge. Check 7 is satisfied by a single comma. Neither is a flaw in the checker: the brief says plainly that judges also read the repo, the docs and the demo video, and that the checker is not the whole assessment. It is a floor. My job, as I saw it, was to find out how far that floor sits from the walls.
The bug that lived for about an hour
The spec has a favourite sentence, paraphrased: hiding another judge's scores in your template is not refusing, because the check has to live in the backend, where curl arrives. I believed it. Before writing real code I wrote the one function that decides who may see judge scores, and put my paper drafts through a stress test. The first draft never checked that the caller was a judge: a participant asking for judge scores computed a target of None, watched None != None come out false, skipped the permission check and got an empty 200, which is exactly what the checker's "participant blocked" question is waiting for. The second draft fixed that so enthusiastically that it locked out organizers too. Both wrong, in different ways. That felt like a triumph, and it was a trap, because afterwards the function was declared authoritative, and authoritative things stop getting looked at. Keep an eye on None. It is about to come back.
It went into the code around hour 25, in the commit that introduced judging. This is the function as it was committed:
# commit d62c0d8, around hour 25
target = judge_external_id or actor.judge_external_id
if target != actor.judge_external_id:
raise PermissionDenied
return Ballot.objects.filter(
assignment__judge__external_id=target,
assignment__event=actor.event,
)
Then the adversarial rounds, which I ran with Gemini and Claude and pointed at the isolation function and nothing else, caught it. One of them asked the question I had not: what about judges who are not in the fixture file?
Fixture judges carry an id like jdg_07. A judge invited through an organizer's link in a real event does not. Their external_id is NULL, and when you hand Django None it writes IS NULL. So every id-less judge matched every other id-less judge, and each could read the others' ballots. About an hour after it landed, this went in:
# commit b410a0a, about an hour later
if judge_external_id and judge_external_id != actor.judge_external_id:
raise PermissionDenied
return Ballot.objects.filter(
assignment__judge=actor.user,
assignment__event=actor.event,
)
Django documents this behaviour openly: an exact lookup with None is interpreted as SQL NULL (docs). It is sensible, and it is exactly the wrong thing to have underneath a security decision. The external_id is still used for one thing, validating an id a caller explicitly asks for. It is never used to decide who the caller is.
The same fix commit repaired two neighbours that the review turned up. Autosave could still edit a ballot after it had been submitted. And scores had no range check at all, so a 9,999 or a minus 3 went straight into the database (the column only limits digits). One review, one commit, three defects. Bugs travel in groups.
Why the checker could never see it
The checker's judges are judge_a and judge_b, both fixture judges with ids, and its peer probe asks for jdg_07, also a fixture judge. The bug only exists in a population the demo data does not contain: real judges, invited the way a real organizer invites them. It would have passed every acceptance run and leaked on the first real event. This is the unglamorous shape of the number-one item on OWASP's list, broken access control: the check exists, and is even correct, for the inputs somebody thought of.
The regression test builds the missing population: two judges with external_id=None on the same project, each with a completed ballot, asserting that judge one sees only their own. It is test_judges_without_external_ids_only_see_their_own_ballots, and it is the one that goes red in the next section.
Five ways to break a portal on purpose
Once I had one bug the checker could not see, the obvious question was how many more kinds there could be. So I wrote a small harness. For each experiment it copies the submitted code to a temp directory, changes one thing, seeds a throwaway database, starts the server, runs the real run.py, and then runs the repo's own tests for that area. This is mutation testing done by hand, at the level of "somebody deleted the security check".
The historic bug (row 2): seven of seven PASS, as promised in the opening, and the regression test goes red alone. The control (row 3): I deleted the peer-score check entirely, and both the checker and the repo test fail. Every experiment needs a control that is supposed to be caught, or you cannot tell a blind checker from a broken harness.
The deadline (row 4) has history. During the build, "closed event refuses submissions" turned green for the wrong reason: the checker's POST was refused by CSRF protection, so the 4xx happened whether or not any deadline logic existed. The fix was precision about who needs CSRF protection. Requests that authenticate with an Authorization: Bearer header skip it, because a browser never attaches that header on its own and a cross-site attacker cannot forge it; cookie-session requests keep full enforcement. Now removing the deadline turns the report to verified nothing, because T1 is the floor.
The CSV (row 5): I replaced the export with placeholder,rows: seven of seven PASS, while the real CSV test (per-project rows, raw and normalized means, ranks, review counts, flags) goes red. The tallies (row 6): showing vote tallies during the voting window is a T3 leak. The checker covers T1 and T2 only, so it says nothing, and two repo tests catch it.
None of this says the checker is bad. One of the three misses is outside anything it claims to check, and the other two slipped through for reasons visible in its own source. A floor tells you the floor holds. "Two of five versus five of five" tells you far more than "seven of seven".
The alarm that stayed for about 45 hours
My favourite piece of boring engineering is the route-by-role matrix. tests/authz_expectations.yaml declares, for every named route, who may call it and what each kind of caller must get back, and a test walks the URLconf and fails the build if any route has no declaration. Forty-seven routes, forty-eight declarations (one extra covers the checker's peer-scores probe).
An alarm nobody has heard go off is a rumour. So around hour 18, in the commit that built the auth and policy layer, I planted a deliberately unprotected route, debug/_leaky_test_only/<judge_id>/scores, and watched the coverage test scream. Satisfying. The docstring said, in capitals, to remove it before the freeze.
The part I am less proud of: it was a plain route in the URLconf, with no debug flag around it. It returned an empty list, so it never actually leaked anything, but it sat in the main branch for around 45 hours, until around hour 63. A backdoor built to test your backdoor detector is still a backdoor, and "I will remove it later" is the most expensive sentence in software.
It finally went in the commit that added the safer drill: a synthetic URLconf holding one undeclared route, and a test asserting the alarm fires and names that exact route.
@override_settings(ROOT_URLCONF='tests.urls_synthetic_undeclared')
def test_coverage_alarm_fails_on_undeclared_route(self):
coverage_test = AuthzCoverageTest()
with self.assertRaises(AssertionError) as ctx:
coverage_test.test_every_route_has_declared_expectation()
self.assertIn('synthetic_control_case_undeclared', str(ctx.exception))
A second test asserts that the production URLconf has no such route and that the old path returns a plain 404. The alarm is still proven, and nothing leaky ships. I should have done it in that order from the start.
What the reviewers found, and what they missed
The adversarial rounds did not stop at isolation. I pointed them at the voting code too, and a few of their findings deserve a table, because each row is a place where my own tests had been quietly agreeing with me.
| Finding | What would have happened | The fix |
|---|---|---|
| Project links were stored raw | A javascript: URL typed into a project link reached an href. Forms validate URLs, but my services bypassed forms. |
Validation in the service layer (absolute http or https, a host, no whitespace or control characters, 1,024 characters at most), plus a render-time safe_link filter so bad data written straight to the database still renders no link. |
| Organizer role was checked against the current event | An organizer of one event could edit a different event. | One helper, require_event_role(actor, event, role), checks the event being edited, with named isolation tests. |
| Voter identity in the audit log | Signed-in votes stored the raw user id, and the audit entry carried actor and project together, so an organizer could connect a voter to a choice. | A secret ballot: an HMAC pseudonym keyed by the server secret and an event seed. Voting audit entries carry no user. |
| Open-link fingerprint included the User-Agent | Cycling the header from one machine counted as a new voter every time. | The fingerprint is the IP address only. |
That last row is easy to check yourself. Five different User-Agent strings from one address, against the submitted code:
UA A/1 -> 302 (accepted)
UA B/2 -> 409
UA C/3 -> 409
UA D/4 -> 409
UA E/5 -> 409
The one nobody found. On the judge's ballot there is a button that opens the project's live site. I wrote the template against a field called live_site_url. The model's field is live_url. A Django template treats a missing attribute as empty, so the button simply never rendered: no error, no log line, no failing test, for every judge on every project. It went in with the same commit as the isolation function, around hour 25, and stayed invisible for well over a day, until the link-handling fix was read line by line. The regression test is called test_judge_ballot_renders_live_url_from_the_model_field.
I would like to say the reviewers should have caught it. A reviewer reads code for what is wrong, though, and a misspelled attribute is not wrong in any way a linter can see. It is merely absent. Absences are what people find, because people open the page.
Green CI, unstyled app
Here is a cousin of that typo, one layer further out. The container runs with DEBUG=false. My settings declared a STATIC_URL, but nothing in a production configuration actually served static files: no STATIC_ROOT, no static-serving middleware, no collectstatic. Every page loads /static/css/rubric.css and /static/js/htmx.min.js. In the container both would have been 404s, which means an unstyled app with every HTMX interaction dead. Django's development server serves static files whenever debug is on, so every local run looked fine, and the acceptance checker only looks at status codes on data routes, so it stayed happy too.
I reproduced it on the commit just before the fix (7632286 is the fix; check out its parent) with DJANGO_DEBUG=false, which is how the container runs. Pages return 200, both static files return 404, and the checker prints 7/7 PASS.
The fix is boring in the best way: WhiteNoise in the middleware, a STATIC_ROOT, and collectstatic as a build step.
RUN cd /app/src && \
DJANGO_SECRET_KEY=build-only-secret \
DJANGO_DEBUG=false \
DATABASE_URL=sqlite:////tmp/collectstatic.sqlite3 \
python manage.py collectstatic --noinput
CI now fetches the CSS and the JavaScript from the running container and fails if either is missing. It landed around hour 53, three quarters of the way in, which is late for a bug that makes your product look like a 1996 homepage.
Proving "runs offline" instead of asserting it. The first rule of the event is that the portal starts with the network off. I did not want to write that sentence, I wanted a job that fails when it is false. The design: build and pull the images while online, add a firewall rule that drops public egress from the Compose network, check that the rule really blocks things (a control request from inside the container must fail and the rule's drop counter must go up), run the acceptance checker, and require the offline report to be identical to the online one. Merely detaching a container from a network proves nothing, because another attached network or a runtime path could still reach outside, so the check watches the actual boundary.
It failed on its first run in CI. The script read the network's ID off a container that had been created but never started, and the ID came back empty. The fix reads the network's name instead. The same commit added a rule that lets already-established connections through, so only new outbound connections are dropped, and a diagnostics step that runs only on failure. That was around hour 60.
The bug you only find by running it
Now the confession. Every check in this post is a machine asking a machine a question. For three days that was the whole development loop, on purpose: I wanted the logic solid and fully tested first, and I told myself the polish, and the looking, would get their turn. I wanted the interface to be clean, quietly sophisticated and easy to use, because a judging portal is a UX problem before it is anything else. I did not run the project by hand until the last day, and by then it was far too late to do anything about what I saw.
So after the freeze I did it properly. I wrote a crawler (a page of Playwright) that logs in as each seeded persona, follows every internal link, and records errors and sideways scrolling at five screen widths.
SIDEWAYS = "document.documentElement.scrollWidth > document.documentElement.clientWidth + 1"
It visited 218 pages at each width, 1,090 page loads in all. (The organizer crawl hit my 70-page cap, so treat that persona as a sample.) The good news is real: no server errors and no tracebacks anywhere. Here is the rest.
| Finding | What happens |
|---|---|
| The organizer's nav does not fit a laptop. | At 1600 px wide it is fine. At 1440 px every organizer page scrolls sideways by 32 px, at 1280 px by 112 px, and "Log Out" slides off the right edge. |
| There is no phone layout. | At 390 px, nearly every page for every persona scrolls sideways, and the nav is cut off after "Results". |
| "My Submissions" is a trap for visitors. | Logged out, the nav link returns a raw JSON authentication_required error. "Create Team" next to it politely redirects to login. Same visitor, two behaviours. |
| The front door had no handle. | See below. |
The front door
This is the one that stings. Run docker compose up, open the portal, and click Log In. Who are you? The seed script prints four logins, but as Authorization: Bearer headers for curl and for the checker. No seeded account has a password. My README, at the submitted commit, told you to sign in as an organizer with a password that did not exist. A judge who wanted to click around my portal had to know how to create an organizer on the command line, and nothing told them.
The funny part is that a working door existed all along. The auth middleware accepts the same token from a cookie named session, so pasting the seeded token into your browser as that cookie logs you in as that persona. I only found it while writing the crawler, because that is how the crawler gets in. It was in none of my docs.
# log in as the seeded organizer in your browser: add a cookie
# name: session
# value: rubric_seed_organizer_tok_9f8e7d6c5b4a
curl -s -o /dev/null -w "%{http_code}\n" \
-H "Cookie: session=rubric_seed_organizer_tok_9f8e7d6c5b4a" http://localhost:8080/organizer
# 200
The root cause is a neat little irony. The checker never logs in, by design. And, as it turns out, neither did my seed script, on behalf of a human. Every layer of evidence I built was shaped like that checker. None of them was shaped like a person arriving at the door. The shape was set on day one: my authentication is a standalone token system, chosen because the checker attaches a header and never logs in.
I fixed the README's login instructions after the freeze, in a docs-only commit, and left the code exactly as submitted. The nav and the phone layout stay as they were. I would rather show you the scar than repaint it.
If you take one habit from this section: put a human-shaped check in by hour 36. Ten minutes of "log in, click everything, shrink the window" would have found every row in that table while there was still time to fix them. It costs a page of code, and mine took a page.
The maths that fought back
The brief asks for cross-judge normalization "documented and defended". The problem in one sentence: some judges are stingy, some are generous, and some give everything the same number, so a raw average partly measures who happened to be assigned to you.
A judge who gave everything a 4
Judge jdg_07 scored three projects and every criterion on every one is a 4. Divide by their spread and you divide by zero, which is a fine way to crash a hackathon.
My method is shrinkage in the spirit of empirical Bayes (deliberately not a full Bayesian posterior, which a statistician would poke). Distrust each judge's own average and spread in proportion to how few ballots you have seen from them, by blending in four imaginary ballots from the average judge. The James-Stein estimator is the famous ancestor.
mu_s = (n * ybar + 4 * mu0) / (n + 4) # shrunk mean
s2_s = (n * s2 + 4 * var0) / (n + 4) # shrunk variance
z = (y - mu_s) / sqrt(s2_s) # ballot value as a z-score
mu0 and var0 are the pool's own mean and variance over all counted ballots. For jdg_07: n = 3, ybar = 4, raw variance 0. The shrunk variance is 0.2246, strictly positive, so the judge gets a finite z-score of 0.5115 and nothing divides by zero. Average a project's z-scores, map back to the 1 to 5 scale, done.
The sentence I had to take back
My first write-up said a constant judge's influence "converges toward zero". It read well, and it was a claim about a formula I had not run on real data. When I ran it, jdg_07 contributed 0.5115 to each of their three projects: a fixed offset that does not differentiate between them. I corrected the docs and logged the correction.
Then the temptation. Remove jdg_07, recompute, and their three projects move 8, 9 and 13 places. Damning? I ran the same experiment on an ordinary judge, jdg_06:
| Judge removed | Project | Rank before | Rank after | Places moved |
|---|---|---|---|---|
jdg_07 (constant) |
prj_09 |
16 | 24 | 8 |
prj_17 |
19 | 28 | 9 | |
prj_19 |
18 | 31 | 13 | |
jdg_06 (ordinary control) |
prj_13 |
31 | 26 | 5 |
prj_26 |
32 | 34 | 2 | |
prj_35 |
26 | 31 | 5 |
The ordinary judge's projects move too; rank is brittle in a tightly bunched field. The biggest jump, prj_19, is the best story in the data: it has only two reviews, one of them jdg_07's flat 4, 4, 4, so its only relative signal is the other reviewer, jdg_29 (3, 2, 5). Take the constant judge away and the project is judged by one person. So I did not assert "constant judges cannot move ranks" in the tests. I wrote the sensitivity down as a limitation.
A second constant judge the brief never mentioned
The spec promises one awkward judge. My dashboard flags two. Judge jdg_19 reviewed four projects, one of which (prj_07) is the earlier half of a duplicate submission, superseded by the team's later prj_41, so its ballot is excluded. The three that remain:
| Project | Functionality, quality, innovation | Ballot mean |
|---|---|---|
prj_03 |
3, 5, 3 | 3.667 |
prj_24 |
3, 4, 4 | 3.667 |
prj_41 |
5, 4, 2 | 3.667 |
Three different score sets, one identical average: zero variance across the ballots the method sees. The flag fires on the weighted ballot value, not on repetitive-looking raw numbers. The duplicate decision moves the headline numbers too: dropping the superseded copy's five ballots changes the pooled mean from 3.5661 to 3.5758 and the variance from 0.4202 to 0.3930. Those two numbers have a story of their own. Around hour 62 I corrected them: an early draft of my judging document said 3.6190 and 0.6019, the real run said otherwise, and prose does not fail a build. Now a test recomputes the statistics from a seeded run and asserts that the document's digits match.
What the pictures show, and what they do not
On the real fixture, 40 projects are ranked from 121 ballots by 29 judges: seven ranks unchanged, fifteen up, eighteen down. A two-way tie for first at 4.3333 raw is broken in favour of prj_34, and the biggest single move is prj_28, down seven places. None of that proves the new ranking is right. A fixture has no ground truth.
The spread of judge averages narrows from 0.3235 to 0.2023, about 37.5 percent. That is largely by construction: shrinkage pulls things toward the middle whether or not the middle is correct. To test closeness to a truth you have to plant one: 20 projects with known quality, 10 judges with leniency offsets from -1.5 to +1.5, a little noise, three judges per project, and a Spearman correlation against the truth.
Over 100 seeds the correlation rises from 0.7347 to 0.8949, and normalization wins in 100 of 100. I did not tune anything, and it is still a simulation whose judges follow the additive model the method assumes, which is why the proof page labels it SYNTHETIC. One more unflattering fact: eight of the 41 projects have fewer than three reviews, and the dashboard says so out loud.
What I planned, what I shipped
My written cut order, made before kickoff, started with pairwise mode and ended with T3 itself. Reality followed it almost to the letter, which is the only reason none of this feels like a disaster.
Most of T4 did not ship: no webhooks, certificates, embeddable widget or OpenAPI spec. A small JSON and CSV surface does exist (judge scores, ballot submission, organizer progress, audit log and verification, a voting summary, comment flagging, the CSV export). The commit history has a small joke about it: django-ninja, picked to generate the OpenAPI schema, sat in requirements.txt from the first hour to the very last, imported by nothing, and left in the commit literally named final.
Pairwise mode and email-gated voting are not built. Comparing two projects at a time and recovering a ranking with a Bradley-Terry style estimator, as Gavel does for HackMIT, is the elegant way around judge calibration, and also a second judging engine. I would rather ship one hard bonus with a simulation behind it than two with excuses. Email-gated voting either fails offline or makes an identity claim I cannot back.
The audit chain has a stated ceiling. Every state change is hash-chained and a database trigger refuses updates and deletes. That makes tampering visible against a trusted head hash. It does not stop someone with full write access from rebuilding the chain, so the honest advice is to publish the head hash somewhere the database cannot reach. "Tamper-evident" and "tamper-proof" are different words, and I only claim the first.
Docs that outlive their folder. My design notes lived in a folder that was always going to be deleted at submission. A late sweep found nearly a hundred references to them in code comments, the stylesheet header and the subtitle of the organizer dashboard. A test now fails the build if any file cites a document that does not ship.
The regret I would fix first is the UI. I planned something clean and easy. What shipped works, has a ballot screen I am happy with, and has a navigation bar that does not fit a laptop. With another day I would not add a feature; I would spend the first hour clicking through it.
If you are building one of these tomorrow
- Ask which population your demo data cannot contain. Fixture data is a tidy little world. My bug lived in the untidy one.
- Never scope access by a column that is allowed to be empty. Scope by the row that logged in.
- Break your own thing on purpose, and keep the drills in the test suite. A check you have never seen fail is a rumour. A deliberately broken route is a debt with a due date, so pay it early.
- Measure before you assert, and write the correction down. "Converges toward zero" felt true and was not.
- Score your checkers. List what each layer cannot see. "Two of five and five of five" is information; "seven of seven" is a mood.
- Run the shipped artifact like a stranger by hour 36. Log in the way a visitor would, click everything, shrink the window, and test the container, not just the dev server. The cheapest bug report in the world is a person arriving at your door.
- Look for absences. A reviewer reads code for what is wrong. A misspelled template attribute is not wrong, it is missing, and it can hide for a day.
The scoreboard
| Thing | Number |
|---|---|
| Acceptance checker | 7 of 7 PASS; T1 and T2 claimed and verified |
| Tests | 271 (270 pass; 1 PostgreSQL-only test skipped on SQLite; the PostgreSQL CI job is green on the evaluated commit) |
| Routes with a declared access policy | 47 (48 declarations) |
| Deliberate breakages caught: checker / repo tests | 2 of 5 / 5 of 5 |
| Isolation leak, committed to fixed | about an hour |
| Hidden live-link typo | well over a day |
| Deliberately leaky debug route, planted to purged | around 45 hours |
| Container before the static fix | pages 200, CSS and JS 404, checker 7 of 7 PASS
|
| Crawl, five widths | 218 pages each, no server errors; organizer nav overflows by 32 px at 1440 wide and 112 px at 1280 |
| Real fixture | 40 of 41 ranked, 121 ballots, 29 judges |
| Judge-mean spread, before to after | 0.3235 to 0.2023 (by construction) |
| Simulated recovery of a planted truth | 0.7347 to 0.8949, better in 100 of 100 seeds (synthetic) |
| Bonuses claimed / not built | Normalization Proof, Threat Model / pairwise mode, email-gated voting, most of T4 |
Run it yourself
git clone https://github.com/Dev-Am12/Rubric.git && cd Rubric
git checkout 777bbb5 # the evaluated commit
docker compose up # seeded portal on http://localhost:8080
python3 run.py .dogfood.toml # the seven questions
pip install -r requirements.txt # once, for the two scripts below
python mutation_check.py . # five breakages, two kinds of evidence
pip install playwright && playwright install chromium
python crawl_check.py http://localhost:8080 # a stranger's walkthrough
Both scripts only touch temporary copies or issue GET requests. I ran them on Linux, and they need nothing beyond the standard library, Playwright and the project's own dependencies. Grab them here: https://gist.github.com/Dev-Am12/9d9782e47cfa9fadff1fd5a6b3f90b9a
The repository is at github.com/Dev-Am12/Rubric. The design decisions, including every correction above, are in DECISIONS.md, and the normalization method is defended in JUDGING.md. Demo: https://drive.google.com/drive/folders/1sVArFhpSnswrgRyhIGcacO7rwGqx6_lW
The line I keep coming back to is small. A green report is a receipt for the questions somebody thought to ask. The useful thing you can do the day after a hackathon is ask a few more, and one of them should be: can a stranger get in?
Worth every hour
I will end on the part no table can hold. I enjoyed this. Seventy-two hours with a shifting calendar, a navigation bar that does not fit a laptop and a checker that was far too polite about my bugs, and I had fun for all of it. Hackathon Raptors events are always a good time to take part in, and DogFood was no exception.
Whatever the result turns out to be, these 72 hours were not a waste. I built something I am proud of, broke it on purpose, learned where my own blind spots live, and wrote it all down so the next person gets to skip a few mistakes. That sounds like a pretty good weekend and a half to me.
References
- DogFood 2026 brief and spec: dogfoodhack.com and dogfoodhack.com/spec
- Rubric, evaluated commit: Dev-Am12/Rubric@777bbb5
- OWASP Top 10 2021, A01 Broken Access Control: owasp.org
- Django
exactlookup andNone: docs.djangoproject.com - Mutation testing: Wikipedia
- Empirical Bayes and the James-Stein estimator: Empirical Bayes method, James-Stein estimator
- Spearman's rank correlation: Wikipedia
- Gavel, HackMIT's pairwise judging system: github.com/anishathalye/gavel
Built solo for DogFood 2026 by Hackathon Raptors. #hackathonraptors #DogfoodHackathon









Top comments (0)