
Picture a judging spreadsheet with one column that never changes. A judge opened the form, gave the first project a 3, and gave every project after it a 3 as well.
DOGFOOD 2026 put a judge like that in the data on purpose. Every team was asked to build the same thing, a submission and judging platform for hackathons, and the organizers shipped a shared fixtures.json with the awkward cases left in: a duplicate submission, review batches nobody finished, and a judge who gave every project the same score.
We built Forgeboard, from first commit to final verdict. This post is about the decisions behind it, and a surprising number of them come back to that one judge.
TL;DR
- Forgeboard is a self-hostable hackathon platform: events, teams, submissions, a public gallery, judge assignment, weighted rubrics, cross-judge normalization and published results.
- One Next.js app, one SQLite file, one container.
docker compose upgives you a seeded portal, and it keeps working with the network switched off.- The official DOGFOOD checker passes 7 of 7. It only has checks for T1 and T2, so T3 and T4 are covered by our own evidence: 131 unit and integration tests, 10 live checks against the running app, and browser suites that drive real Chrome.
- Code (MIT): https://github.com/Avi36005/DogFood-2026
What we built
Four kinds of people use a hackathon platform, and each of them gets a different product:
- Organizers set up tracks, prizes, custom questions and deadlines that the server enforces, then run judging from a console.
- Teams join by invite link, save drafts and submit before the deadline.
- Judges score against a weighted rubric and only ever see their own assignments.
- The public browses the gallery, votes, comments and reads the published results.
| Tier | What's in it |
|---|---|
| T1 · Core | Accounts with local password recovery, per-event roles, event setup (dates, timezone, tracks, prizes, five kinds of custom question), invite-link teams, draft-and-edit submissions, server-enforced deadlines, a searchable public gallery |
| T2 · Judging | Scoped judge invitations, previewed batch assignment, organizer-weighted rubrics, backend-enforced isolation, eligibility decisions, live progress, documented normalization, CSV at every stage |
| T3 · Public | Community voting in three access modes, per-voter randomized ballots, totals hidden until published, rate limits, moderated comments |
| T4 · Stretch | REST API (14 operations, OpenAPI 3.1, scoped keys), signed webhooks, certificates, Ed25519-signed judge records anyone can verify, an embeddable gallery, whole-event import and export |
The organizer's workspace opens on what is waiting on them, not on a wall of charts: an unassigned project, a judge who hasn't started, results ready to review.
Rule 1: a stranger, a laptop, and no network
The spec's first hard requirement is that docker compose up brings up a working, seeded portal with the network off. We treated that as the design constraint rather than a deployment detail. Our architecture doc puts it bluntly: the thing being optimized is whether a stranger can run it, and every additional moving part is a way for that to fail on someone else's laptop.
So the whole system is this:
browser ──HTTP──▶ Next.js (App Router, React 19)
pages and route handlers ← no rules live here
capabilityFor() ← the only door
domain services ← every rule lives here
node:sqlite ← one file on a volume
- No second process. No database container, no queue, no cache, no sidecar.
-
SQLite through
node:sqlite, which ships inside Node 24: no native build step, no driver to install, nothing to start. - No ORM. The schema is plain SQL in one file that a reviewer can read in a single pass.
-
Nothing phones home. Fonts are vendored from the
geistpackage, analytics are gone, and apart from webhooks you configure yourself, the app makes no outbound requests at runtime. - No mail server. Team invites, judge invitations, voting tokens and password recovery are all one-use links that an organizer or admin hands over.
That bought us a cold start to a passing healthcheck in 1.2 to 1.6 seconds in our runs, a container started with --network none that still serves every public route, and restarts that keep the data and never re-seed.
The fixture import follows the same instinct: boring, and loud about anything it changes.
[forgeboard] imported evt_01 from fixtures.json: 40 projects, 40 teams, 30 judges, 121 scores
[forgeboard] duplicate submission: prj_07 (replaced by prj_41); 5 of its reviews not imported
[forgeboard] seeded. test logins (published in .dogfood.toml; demo only):
organizer Cookie: forgeboard_session=org_demo_7f2a9c41d8e3b6a5
judge_a Cookie: forgeboard_session=jdg_a_demo_91bc5e0f27d4a8c3
...
Our schema allows one project per team, enforced by a unique index, so the importer keeps the resubmission and says out loud what it dropped. Quietly "fixing" data is how you end up with a ranking nobody can explain.
The shared DOGFOOD fixtures, imported at boot and live in the public gallery. Dry Harbour, the resubmission, is the copy that stays.
Rule 2: hiding a button is not refusing a request
The DOGFOOD spec tells the story of a team whose checker run fails on a single line: judge cannot see peer scores ... got 200, wanted 401 or 403. Their template hid the other judge's scores, and their API returned them to anyone logged in. The spec's verdict: "That is the difference between hiding a button and refusing a request."
We built the product around that sentence.
Rules live in the data layer. On every request, Forgeboard builds a capability, meaning what one actor may do inside one event, from the event_roles table. It is never cached between requests and never sent to the browser. Domain functions refuse to run without the right one:
export function judgeProgress(cap: Capability): JudgeProgress[] {
requireOrganizer(cap); // throws AccessDenied
// ...
}
Isolation is a query shape, not a filter. A judge's queue isn't "every assignment, filtered down to mine". It is selected by ownership, so there is no code path that could return someone else's:
SELECT ... FROM assignments a
WHERE a.event_id = ? AND a.judge_user_id = ? -- the actor's own id
Opening a single review re-runs every check instead of trusting that the judge arrived from their queue: the assignment belongs to this event and to this judge, the project sits inside the judge's track grant, and the judge is not, and never has been, on the project's team. We keep team membership history, so leaving a team after assignment doesn't clear the conflict. Saving a review calls the open path first, so a write can't skip the read's checks. And the API routes call the same domain functions as the pages, so curl is refused exactly the way the interface is.


Sam, a judge, pastes another judge's review URL and gets a plain 404. The organizer's audit trail has already recorded it: review.open, denied, "not the assigned judge".
Why a 404 and not a 403? A judge poking at an organizer URL learns nothing, not even that the page exists. The machine-facing endpoints, CSV exports and the API, answer 403 instead, because there a clear refusal is more useful than a polite fiction.
The bug we designed around. In the Next.js App Router, a route-level
loading.tsxwraps the page in a Suspense boundary and flushes the shell immediately. Once a byte is on the wire,notFound()andredirect()can no longer set a status code, so an unauthorized page would answer 200 with a client-side redirect inside it. Status codes are part of our authorization story, so Forgeboard has no route-level loading files. Where a skeleton genuinely helps, like the gallery grid, the Suspense boundary sits inside the page, after the access decision has been made.
One more detail we like: an instance admin can open any event, because that's what administration means. It just isn't quiet. The first touch writes an admin.access row into that event's own audit trail, where its organizers will see it.
Rule 3: never invent an opinion a judge didn't give
Scoring a single review is the easy part. The organizer defines criteria, each with a weight and an integer scale, and a review's raw score is the weighted mean:
raw = Σ(wᵢ · sᵢ) / Σ(wᵢ)
It's computed at submit time and stored, so editing the rubric later can't retroactively change a score someone already gave. Our demo rubric borrows DOGFOOD's own scoring weights: Completeness 40, Integrity 25, Operability 20, Craft 15.

A judge scores Beacon 4, 5, 3 and 4 against weights of 40/25/20/15. The running total reads 4.05, the same number our unit tests pin (76.25 out of 100).
The hard part is that raw scores aren't comparable across judges. From our seeded event:
| Judge | Reviews | Raw mean | Bias vs. panel |
|---|---|---|---|
| Mo Nakamura | 5 | 2.06 | −1.03σ |
| Lena Batista | 5 | 2.55 | −0.56σ |
| Panel | 3.12 | 0 | |
| Yuki Costa | 5 | 3.91 | +0.76σ |
| Jai Batista | 4 | 4.31 | +1.15σ |
A project that draws Mo and a project that draws Jai are not being measured with the same instrument. Average the raw scores and the panel lottery decides part of the ranking.
So Forgeboard standardizes within each judge:
z_jp = (raw_jp − mean_j) / sd_j # population sd, per judge
score_p = mean of z over the project's usable reviews
display = clamp(panel_mean + score_p · panel_sd, scale_min, scale_max)
Then you meet the judge who gave everything a 3. Their standard deviation is zero, so the formula divides by zero. The textbook escapes are to borrow the panel's variance or to shrink the judge toward it. Both manufacture separation from a ballot that contains none. A judge who gives every project a 3 has told you nothing about which project is better, and the honest response is to say so rather than invent an opinion on their behalf.
Forgeboard only standardizes a judge with at least three completed reviews and a spread above zero. Anyone else is excluded from the normalized calculation, by name, with the reason, and their raw scores stay on screen. A project left with fewer than two usable reviews gets no rank at all: it's labelled insufficient comparable reviews instead of being handed a confident position. If every judge were excluded, the API, the console and the snapshot would all say raw_fallback. The method never switches silently.

Judge calibration, before anything is published: every judge's raw mean, spread and bias. Mo Batista is excluded for no variation, Umi Costa for too few reviews, and both are named.
On the seeded event it changes a lot. 32 of 36 ranked projects move. Windlass is first either way. Notary climbs from raw #11 to #4. Trellis slips from #3 to #5, because it drew a generous panel. Vantage falls furthest, from #6 to #17.

The first 17 of 36 rows of a snapshot. Normalized score, raw score, raw rank and movement sit side by side, so the adjustment is visible instead of asserted.
Snapshots are immutable and record the method, its version and its parameters, so a published ranking carries a record of how it was made. Computing one changes nothing public. Publishing is a separate, confirmed step, because it is the one action in the product the whole world can see.

The one confirmation in the product, in plain words: publishing makes the standings public to everyone.
You don't have to take our word for any of it. normalize() is a pure function with no database behind it. Export reviews.csv and criterion-scores.csv and you can recompute a published ranking yourself. The tests pin the arithmetic down to [2, 3, 4] standardizing to [−1.224744871, 0, +1.224744871].
Fair before anyone scores
Normalization corrects for judges. Assignment decides who judges what in the first place, and it's easy to get quietly wrong.
-
Order isn't submission order. Projects are ordered by a SHA-256 of
seed | projectId, so the first team to submit doesn't get a systematically different panel from the last. - Lightest load first, among judges who are eligible: inside the project's track, and never on its team, past or present.
- Additive, not destructive. Existing assignments are counted and kept, so running it again after a late batch of submissions tops up the gaps instead of reshuffling panels that have already started.
-
Honest about failure. A project that can't be fully covered comes back in a
skippedlist with the reason, instead of being silently dropped. -
Preview is a real dry run. The console runs the same planner with
dryRunset and writes nothing. Committing runs it again on the server, so what comes back from the browser is a decision to proceed, never the plan itself.

Preview shows the load per judge after the batch and the projects still short. "Nothing has been written yet."
Testing like we didn't trust ourselves
Our committed acceptance-report.txt ends with this line:
claimed T1 T2 T3 T4, verified T1 T2
That's the correct output. run.py only has checks for T1 and T2, and the spec is clear that honest reporting beats a README claiming everything works. So we left the report exactly as the checker printed it and wrote our own evidence for the rest:
- 131 unit and integration tests, covering the scoring maths, isolation, deadlines, voting, signing and tamper rejection, import validation and governance.
- A 175-line headless Chrome driver that speaks the DevTools Protocol over Node's built-in WebSocket. No Playwright, and nothing to install. It drives the whole lifecycle in a real browser: register, create and configure an event, invite a judge, form a team, fail a submission, fix it, submit, judge, disqualify, reinstate and recover a password.
- axe-core on 15 screens with zero serious or critical violations, plus keyboard and focus checks and a width sweep at 320, 390, 720, 1024 and 1440 pixels with no horizontal overflow.
- A live T4 script: 10 of 10 checks pass against the running instance.
- The authorization matrix over real HTTP, with real session cookies: organizer, judge, participant and anonymous requests against every sensitive route, and the refused attempts on the audit trail.
And a seed that is messy on purpose: one judge who scores everything a 3, one who never started, one who left half a batch in draft, one with only two reviews, and one submitted project with no assignments at all. If your demo data is tidy, you never see the screens that judging is actually about.
What we didn't solve
Writing this list was part of the work:
- No pairwise judging. Bradley–Terry sidesteps calibration entirely by never asking for an absolute score, and it's the better answer to this problem. It isn't in Forgeboard.
- No inter-rater reliability. We show each judge's bias and spread, but not how much judges agree on the same project.
- The thresholds are defaults, not tuned constants. Three reviews per judge and two usable reviews per project are recorded on every snapshot, but not yet editable in the interface.
- Standardization assumes comparable slices. If an organizer hand-assigns every strong project to one judge, that judge's high mean reads as generosity and gets corrected away.
- Ties break by raw mean, then project id. Deterministic, but arbitrary. With genuinely tied scores, an organizer should decide.
- Sybil voting isn't stopped. Our threat model says so and explains why.
Five things we'd tell the next team
- Write the failing check before the feature. If no code path returns another judge's scores, there's nothing to hide.
- Treat status codes as part of authorization, and check what your framework's streaming does to them.
- Seed the awkward cases. The constant judge, the unfinished batch and the uncovered project are the screens judging is actually about.
- When the data can't support a confident answer, say so in the product. "Insufficient comparable reviews" beats a confident #7.
- Count your moving parts. Each one is a way to fail on someone else's laptop.
Try it
git clone https://github.com/Avi36005/DogFood-2026
cd DogFood-2026
docker compose up # seeded and listening on http://localhost:3000
python3 run.py .dogfood.toml # 7 of 7 PASS
Every demo account uses the password forgeboard2026: organizer@, judge@, participant@ and admin@forgeboard.local.

The final verdict: published standings with the raw mean beside every normalized score, and a link to how it was calculated.
The DOGFOOD spec signs off with "Build the boring parts well." Forgeboard is mostly boring parts: a query that can only select your own rows, a status code that stays honest, and a judge excluded by name instead of quietly averaged in. We think that's where a judging platform earns trust.
Team Kryptonite: Avinash Gehi, Hardik Hinduja, Sahil Deshmukh


Top comments (0)