DEV Community

Cover image for I Built the Platform That Judges Hackathons in 72 Hours. The Hardest Bugs Were About Evidence, Not Features.
Rajpriyan S
Rajpriyan S

Posted on

I Built the Platform That Judges Hackathons in 72 Hours. The Hardest Bugs Were About Evidence, Not Features.

A hackathon platform has one job that actually matters. When it says "this project came 14th, and here is why," that sentence has to still be true next week.

That sounds obvious. It is not. I built Verdict Ledger for DOGFOOD 2026, the Hackathon Raptors event where everyone builds the same submission-and-judging portal and the organisers plan to self-host the winner. Most of my hard problems were not features. They were about evidence: scores that change after a ranking is computed, explanations that quietly rewrite themselves, audit logs that fail to chain, and judge data that leaks through a page nobody thought to protect.

This is the write-up of what actually broke, what I built to stop it, and where my own claims stop.

The short version

  • Claimed T1 + T2. Official checker: 7/7 PASS. T3 and T4 deliberately not claimed.
  • 273 tests pass, none skipped, in about 90 seconds, with only the standard library's unittest.
  • A normalization method I can defend in a table: it moves 32 of 41 fixture projects, and one project jumps 17 places because a single judge's lone score was neutralised.
  • Publishing a ranking built on stale scores returns HTTP 409. Explanations read from an immutable snapshot.
  • Every bug below was found by an audit pass over my own build, and each one has a regression test that I watched fail before trusting it.

Repo: github.com/rajpriyanid-creator/dogfood


The brief looks like CRUD. It is not.

On paper the product is simple: auth, teams, submissions, a gallery, a judge console. But the brief describes an event as a ten-stage pipeline, each stage with its own state feeding the next, and it names the stage where things break: normalization.

Two sentences in the spec set the real bar:

A judge must never see another judge's scores. If this only works because the UI hides a button, it does not work.

"We averaged the scores" is an answer, and it is a weak one. Tell us what you did about the judge who marks everything a 3.

Everything below is my attempt to answer those two sentences honestly.


Decision zero: pick a tier and tell the truth about it

I claimed T1 and T2 and wrote, in the README, that T3 (voting, comments) and T4 (REST API, certificates) are out of scope. The scoring says a clean T2 beats a broken T4, and overclaiming costs more than the tier was worth.

The part I care most about is how the README reports status. Every feature gets three columns instead of one tick:

Area Implemented Tested Verified by official checker
Judge cannot read peer scores yes yes yes (1 check)
Deadline enforced server-side yes yes yes (1 check)
Hash-chained audit log yes yes not covered
Normalization freshness gate yes yes not covered
T3 (voting, comments) no - -

"Implemented" and "verified" are different claims. The official checker exercises only seven behaviours. Most of what I built, including what I am proudest of, is not covered by it, and the README says so out loud. The votes and comments tables exist in the schema as a starting point, but nothing reads or writes them, and SECURITY.md states that "this feature doesn't exist yet" instead of describing an attack surface that is not there.


Architecture: one process, one file

One Flask process serves server-rendered HTML for humans and a JSON API under /api for the checker. One SQLite file is the entire persistence layer. No queue, no cache, no frontend build step, no client-side JavaScript.

browser / curl ──▶ Flask (app.py: routes)
                      │
                      ├─ core.py           identity + role-guard decorators
                      ├─ auth.py           sessions, PBKDF2 passwords, token hashing
                      ├─ events.py         the one timestamp → state mapping
                      ├─ assignment.py     deterministic judge assignment
                      ├─ normalization.py  pure maths, no DB dependency
                      ├─ audit.py          hash-chained log
                      └─ db.py ──▶ SQLite (named Docker volume)
Enter fullscreen mode Exit fullscreen mode

The data flows in one strict direction, and I wrote it down because every bug below is a break in this chain:

AUTHORIZATION → SUBMISSION → JUDGING → NORMALIZATION
              → HISTORICAL EVIDENCE → FRESHNESS CHECK → PUBLISH
Enter fullscreen mode Exit fullscreen mode

The lifecycle is one function

Deadline bugs happen when every route does its own date comparison. So events.py maps timestamps to exactly one state, and every permission is a lookup against it:

UPCOMING → OPEN → SUBMISSIONS_CLOSED → JUDGING → PUBLISHED   (ARCHIVED: sticky)
Enter fullscreen mode Exit fullscreen mode
Action Allowed while
create/edit/submit project, create/join team OPEN
score a project SUBMISSIONS_CLOSED or JUDGING
run normalization SUBMISSIONS_CLOSED or JUDGING
publish past OPEN, and a normalization run exists

The status is always computed from the server's clock, never from a client flag. A test sends spoofed event_status and submissions_close fields and confirms they are ignored.


Role isolation: if I can curl it, it is broken

The brief's isolation matrix is blunt: a judge reads their own scores and nothing else. I decided the permission model explicitly instead of inheriting defaults:

Role Can do
visitor browse the gallery and submitted projects
participant create/join a team, create/edit/submit projects for their own team
judge read own scores, review projects assigned to them
organizer everything under /organizer, CSV export
admin organizer powers. Not a participant or judge.

The admin decision. It is tempting to make admin a superuser. I made admin organiser powers only. An account that can both administer an event and score projects in it is exactly the conflict of interest a judging platform exists to prevent. A test asserts that admin gets 403 on /team, /projects/new, /judge and /api/judge/scores.

Identity comes from one place. core.py::current_identity resolves the requester from their session token and nothing else. No code path reads a role, user id, judge id or team id from the query string, form body, JSON body or header and trusts it. Where a request names an id, such as ?judge=jdg_01, it is either ignored or compared against the session and refused on mismatch.

You can check this in 30 seconds. Log in as judge2 and ask for judge 1's scores:

GET /api/judge/scores?judge=jdg_01      →  403 Forbidden
GET /api/judges/jdg_01/scores           →  403 Forbidden   (path form, same rule)
Enter fullscreen mode Exit fullscreen mode

The test that enforces it: 43 declared policies

Rules are only as good as the next route someone adds. tests/test_access_matrix.py declares a policy for 43 method-and-route combinations, then hits every one of them as every identity (anonymous, participant, judge, organizer, admin) and fails if any combination behaves differently from its declared policy. Forget to declare a new route and the build breaks.

The part I got wrong at first is the oracle. A naive version asserts "refused means 401 or 403." But plenty of correct refusals share those codes: a bad password returns 401, a closed event returns 403, a judge not assigned to a project returns 403. None of those are role decisions, so a test that accepts any 403 can pass for the wrong reason. The final test only accepts the role guard's own exact refusal bodies:

guard_refused = (resp.status_code == 401 and body.get("error") == "authentication required") \
    or (resp.status_code == 403 and body.get("error") == "forbidden")
should_be_refused = who not in allowed
if guard_refused != should_be_refused:
    failures.append(f"{method} {rule} as {who}: got {resp.status_code} ...")
Enter fullscreen mode Exit fullscreen mode

A test that passed for the wrong reason

A second, nastier version of the same lesson. My cross-team test claimed to prove that a participant cannot submit on behalf of another team. It passed. But it only passed because the fixture event happens to be closed, so the request failed on the deadline, never reaching the team-spoof condition at all.

Once draft creation stopped being deadline-gated, I rewrote the test to assert the real claim directly: send a hostile team_id: "live_tm_2" and check that the resulting project is attached to the caller's own team. A green test that is green for an unrelated reason is worse than no test, because it tells you the door is locked when you never tried the handle.


The judge who marks everything a 3

The official fixture is a closed historical event: 41 projects, 40 teams, 30 judges, 8 tracks and 126 scores, deliberately built with the edge cases every evaluation system must survive. I ran my normalization against it and wrote down what the data actually looks like.

Fact about the fixture Value
Global mean raw score (demo weights) 3.568
Global standard deviation 0.671
Judge averages from 2.000 to 4.250
Judges with 2 or fewer reviews 8 of 30
Zero-variance judges 3: jdg_01 (n=1), jdg_23 (n=1), jdg_07 (n=3, all 4.0)

A judge average range of 2.25 points on a 5-point scale is the whole problem in one number. Averaging raw scores lets judge severity decide the podium.

The method, stated plainly

It is a transparent, sample-size-shrunk location-scale heuristic inspired by empirical-Bayes reasoning. It is explicitly not full empirical Bayes: there is no hierarchical model and no fitted prior, and k is fixed, not estimated. I refuse to call it more than it is, in the code and in the docs.

For each review, S = Σ(wᵢ × xᵢ). Then, with k = 3 as a fixed design parameter, each judge's mean and variance are shrunk toward the global ones:

μ*_j = (n_j·μ_j + k·μ_g) / (n_j + k)
v*_j = (n_j·v_j + k·v_g) / (n_j + k)

z = (S − μ*_j) / σ*_j
N = clip(μ_g + z·σ_g, 1, 5)
Enter fullscreen mode Exit fullscreen mode

and for a judge whose observed variance is exactly zero:

N = μ_g        # no relative signal, so the review carries no discrimination
Enter fullscreen mode Exit fullscreen mode

The code is the formula, with no hidden state, which is why a rerun is reproducible by construction:

def normalize_one(raw_score, judge_stats, global_stats, lo=1.0, hi=5.0):
    if judge_stats.zero_variance:
        return global_stats.mean, None
    if judge_stats.shrunk_sd == 0:
        # not reachable with k > 0, but guarded rather than dividing by zero
        return global_stats.mean, None
    z = (raw_score - judge_stats.shrunk_mean) / judge_stats.shrunk_sd
    n = global_stats.mean + z * global_stats.sd
    return max(lo, min(hi, n)), z
Enter fullscreen mode Exit fullscreen mode

Why k = 3? It is a chosen constant, not derived from the data. A judge with 5 or more reviews, the bulk of the fixture, is only lightly shrunk, while a judge with 1 or 2 reviews is pulled hard toward the global distribution. A couple of data points should not be trusted like a dozen.

Two rules I made on purpose: missing reviews are never imputed as zero (a judge's statistics use only reviews they actually submitted), and n = 1 is not a special case. A single review has zero variance by definition, so it falls naturally into the zero-variance branch.

What it actually does to the ranking

I ran the real run_normalization over the real fixture with the demo weights (40% functionality, 35% quality, 25% innovation, which is my own demo choice, not something the fixture specifies):

Measure Result
Projects whose rank changed 32 of 41
Projects that moved 3 or more places 13
Rank correlation, raw vs normalized (Spearman) 0.927
Top 5 overlap 4 of 5

The top of the table barely changes (Iron Switch stays first), which is reassuring. The middle is where judge severity was doing damage.

The biggest riser: "Dry Harbour" (prj_07), rank 31 → 14. Its raw average was 3.340 and its normalized average is 3.667. The reason is one review. jdg_01 is a judge with a single review in the whole event, and that review was a harsh 2.0 for this project. One data point cannot say whether a judge is strict or this project is weak, so it is pulled back to the global mean (3.57). The other four reviewers scored it 2.35, 3.65, 4.6 and 4.1.

The biggest faller: "Small Relay" (prj_19), rank 17 → 30. Its two reviews were 3.15 from jdg_29 and a flat 4.0 from jdg_07. That 4.0 is one of three identical scores from a judge who rated everything the same, so it is also neutralised, to 3.57. The project loses the benefit of a generous score that carried no information.

I am telling you about the second one on purpose. The method is doing what it says, but a 13-place drop from neutralising one flat 4.0 is exactly the kind of effect a skeptical organiser should be able to see and argue with. That is why the explain page shows the raw score, the normalized score, and each judge's shrinkage.

The limits, written down

JUDGING.md has a "Limitations of the method" section, and I stand by every line:

  • k = 3 is fixed, not tuned. A different event might warrant a different value.
  • The zero-variance rule is a bright line. A judge giving 4, 4, 4 hits the fallback, while 4, 4, 4.5 does not. It is exact equality, not a threshold.
  • No cross-track calibration. Normalization runs across the whole event, so a systematically stricter track is not adjusted separately.
  • It is not a peer-reviewed method. It is a documented, transparent heuristic, not validated against ground truth, and I claim no such validation.

The bugs an audit pass found in my own build

After the first working version, an audit pass over my own code found real defects. I kept every one visible in the test suite as a labelled regression test, because the story of what broke is more useful than a clean-looking repo.

1. The audit log that could not verify itself

The hash-chained log was supposed to be tamper-evident. In the first version, the seq column was never populated, so every row had seq = NULL, ORDER BY seq ordered nothing in particular, and verify_chain() returned False as soon as a second record existed. A tamper-evidence feature that reports tampering on a perfectly healthy log is worse than none.

The fix made seq an autoincrementing rowid alias. There was a second problem hiding behind it: two concurrent writers could both read "the last row" before either inserted, and both chain to the same predecessor. So the append now takes SQLite's write lock before the read:

conn.execute("BEGIN IMMEDIATE")      # acquire the write lock up front
try:
    last = conn.execute("SELECT hash FROM audit_events ORDER BY seq DESC LIMIT 1").fetchone()
    prev_hash = last["hash"] if last else "genesis"
    row_hash = _hash_row(prev_hash, at, actor_id, actor_role, action, resource, result, detail_json)
    conn.execute("INSERT INTO audit_events (...) VALUES (...)", (...))
    conn.commit()
except Exception:
    conn.rollback()
    raise
Enter fullscreen mode Exit fullscreen mode

2. The integrity dashboard that collapsed 30 judges into one row

The organiser integrity page used a bare COUNT(*) with no GROUP BY, so the whole event collapsed into one row. A second flaw sat behind it: the query joined on assignments, so a score with no assignment row simply vanished.

The regression tests parse the rendered per-judge table and pin exact values, and the expected numbers are recomputed independently from fixtures.json, not copied from the app's output, so a collapsed or arbitrary result cannot pass just by containing the right judge ids somewhere on the page. When I re-added the bad query to check, 11 of the 19 integrity tests failed.

3. Drafts were public by predictable id

/projects/<id> had no visibility check at all, so a draft project's detail page was fully public to anyone who guessed an id like prj_42. The fix returns 404, not 403, unless the project is submitted, or the caller is on the owning team, or is an organiser. A 403 would confirm the project exists. Drafts are also excluded from the gallery and from other projects' duplicate-candidate lists.

4. A score that changed weights but kept its old rubric version

Rubrics are versioned, and each score pins the version it was given under. In an early version, updating a score with a newer rubric left the old rubric_version_id in place while recomputing raw_weighted against the new weights. The stored record claimed one rubric and used another.

The key regression test publishes rubric v2 with wildly different weights (5% / 5% / 90%), then re-runs normalization and asserts that every existing score's raw value is unchanged, because each score must still be computed with its own version's weights. If normalization used the newest rubric for everything, every raw score would change. When two versions are mixed in one result, the result's rubric_version_id is NULL rather than picking a proxy, and normalization refuses incompatible score ranges instead of pooling incomparable values.

5. Scoring that stayed open forever

There was no server-side judging-state enforcement at all, so scores could be created or edited regardless of the event's state, including after results were published. That is where the single-function lifecycle above came from. Once judging starts, tracks, prizes, judges, rubric and assignments are also frozen.

6. A half-seeded database that looked fully seeded

The seed functions each committed their own work, so a failure partway through left a partially seeded database that the old is_seeded() check ("any events exist?") would mistake for complete on the next boot. The seed is now one transaction with a completion marker written last, and a test forces a mid-seed failure to prove it.

7. Sessions that never ended

Once created, a session authenticated forever, and logout only cleared the browser cookie. The same token presented through a raw header kept working. Sessions now expire after 24 hours, and logout sets revoked_at on the server-side row so the token is rejected whether it arrives as a cookie or a bearer header. Revoking one session does not touch anyone else's.


Evidence that cannot move: the freshness gate and the snapshot

This is the idea the project is built around.

Failure mode 1: publishing a ranking built on scores that no longer exist. The flow is normalize → inspect → publish. If a score is edited in the gap, the stored ranking was computed from data that is gone. When a run is created, I store a SHA-256 fingerprint of the judging state, and publish recomputes it. A mismatch returns HTTP 409 and the organiser must re-normalize.

The fingerprint covers every result-affecting input: rubric versions and their criteria (key, weight, min, max), judge assignments, score records, and criterion-level values, all sorted deterministically. It deliberately does not include timestamps or comments. The docstring states the rule I held myself to: it "must exactly match the normalization engine's actual inputs," so that a change to any result-affecting input changes the fingerprint, and a change to a non-result field does not cause a false alarm.

The regression test is the demo flow: normalize, then edit one criterion value with value = value + 0.1, then publish and assert a 409 whose message mentions "stale", then assert that no publication happened.

Failure mode 2: an explanation that rewrites itself. If the "why did this change?" page read live score_criteria, then normalizing, publishing and then editing a score would silently change the explanation of a historical result. So when a run is created, the exact criterion values it consumed are copied into normalization_run_criteria, and the review comment into review_normalizations. explain_result() reads only from those snapshots. The snapshot is write-once, and a new run creates its own.

The two mechanisms are a pair. The fingerprint stops a new result being built on moved evidence, and the snapshot stops an old result's evidence being rewritten.


A hash-chained audit log, and an honest sentence about its limits

Every write action, including denied access attempts, gets a row whose hash covers its own fields plus the previous row's hash. audit.verify_chain() recomputes it from "genesis".

The README calls it tamper-evident, not tamper-proof, and SECURITY.md says why. Someone with write access to the SQLite file can rewrite the whole table, recomputing every hash correctly, and the chain will verify as intact. True tamper-proofing would need an external anchor, such as publishing the chain's tip hash somewhere the operator does not control. I did not build that, so I do not claim it. The chain also only covers what passes through audit.record(), so a direct SQL edit to a scores row is not logged and not detected.


Small security decisions that cost almost nothing

  • Session tokens are hashed. The raw token (secrets.token_hex(20), 160 bits) is returned once. The database stores only SHA-256(token) as the primary key, so a leaked database does not leak live sessions. Passwords use PBKDF2-HMAC-SHA256 with 120,000 iterations and a per-user salt.
  • repo_url is parsed, not pattern-matched. Only http and https with a hostname are accepted, case-insensitively, on both create and edit. javascript:, file: and data: return 400.
  • Rate limiting covers login, team-join and judge-invite and returns 429 with Retry-After. It is in memory and process-local, so it resets on restart and is not shared across workers. The docs say so.
  • Judge "invitation" is direct provisioning. There is no mail service offline, so the organiser creates the judge and a random password is shown once. There is no password-reset flow, by design.
  • Invite codes are secrets.token_hex(6), 48 bits of entropy.

The one I would flag to any organiser: the fixed demo tokens (org_demo_token, jdg_a_demo_token and friends) are published in the repo and never expire, so .dogfood.toml stays valid across restarts. Anyone who has read the repo has organiser access to a default install. SECURITY.md says in bold to delete or rotate them before running a real event.


Offline means really offline

The rule is that docker compose up produces a working, seeded portal with the network off.

Vendored wheels. Flask and its six transitive dependencies sit in vendor/ as 7 wheels, and the Dockerfile installs with pip install --no-index --find-links=/app/vendor. There is no <script src="https://..."> anywhere in a template and no web fonts, only system stacks.

An atomic, idempotent seed, as described in bug 6.

A two-event model. evt_01 is the supplied fixture, loaded verbatim and kept closed and published. Its dates are never altered, because the checker's "closed event refuses submissions" test depends on that date already being in the past. evt_live_2026 is an open demo event with its own teams, judges, rubric and assignments, so the whole create → submit → assign → judge → normalize → publish flow can run live without touching historical evidence.

Two fixture details are handled as evidence, not errors. The duplicate "Dry Harbour" pair (prj_07 and prj_41) loads as two independent rows and each project's page surfaces a "potential duplicate submission" callout. And because the fixture only records outcomes, not the assignment process, seeding reconstructs one assignments row per fixture score, while the assignment engine itself is exercised on the live event.


How I know my tests are not lying

A passing test proves little until it has been seen to fail. For the audited defects and the security boundaries, I temporarily reintroduced the bug or weakened a guard and confirmed the suite went red, then restored the code. Re-adding the bare COUNT(*) failed 11 of 19 integrity tests. Removing a role guard, widening a role, or adding an unguarded route each fail test_access_matrix.py. Using the newest rubric's weights for every score fails the rubric-mixing regression.

This is a manual practice, not an automated mutation-testing tool, and I did not apply it to every test. The README says exactly that.

The numbers at the end of the build:

python3 -m unittest discover -s tests -p "test_*.py"
Ran 273 tests in 87.364s
OK

python3 run.py .dogfood.toml
claimed T1 T2, verified T1 T2   (7/7 PASS)
Enter fullscreen mode Exit fullscreen mode

The official checker's receipt is committed as acceptance-report.txt, generated from a clean docker compose up --build -d with the byte-identical supplied run.py.


What I cut, and what I do not regret

  • Client-side JavaScript. Every action is a full page load. With no client logic there is no client logic to trust, so the server is the only place a rule can live.
  • T3 and T4. I would rather say "this feature doesn't exist yet" than ship half-working endpoints.
  • Email invitations and password reset. No mail service offline, so I did not fake one.
  • TLS. Plain HTTP only. It belongs behind a TLS-terminating proxy for any real deployment.
  • Collusion detection. SECURITY.md says "None, and none is claimed." Two colluding judges look identical, from score values alone, to two judges who genuinely agree. The audit log supports a human investigation, and that is all.

The decision I would take back

My audit log has a gap I documented and did not close. Most business writes commit first and append their audit row afterwards. A crash in that narrow window can leave a mutation with no audit row. The hash chain itself is atomic, but the pairing of "change the data" and "record that I changed it" is not one transaction.

Closing it needs a larger refactor so that both writes share one transaction, and I ran out of window before I ran out of plan. For a platform whose whole thesis is that evidence cannot quietly go missing, that is the honest weak spot, and I would fix it first.

The second gap I think about is result replay. A normalization run stores everything needed to recompute it, and re-running gives identical numbers, which is tested. But there is no button that recomputes a published result and reports a match. A skeptical organiser still has to take my word, or my tests' word, that it reproduces.


What I would tell someone building this next

Decide what your evidence is before you decide what your features are. The brief reads like a CRUD app. The product is a set of claims, and each claim needs evidence that cannot move.

Check that your test is red for the right reason. Two of my worst bugs hid behind tests that passed for an unrelated cause: a team-spoof test that only failed on the deadline, and a role test that accepted any 403. Make the oracle as specific as the claim.

Make forgetting a failing test. Forty-three declared policies, enforced against every role, mean a new route cannot ship without a decision.

Name your limits in writing. A fixed k, a tamper-evident log, an in-memory rate limiter, a crash gap in the audit trail. Each costs one sentence in the README, and having a reader discover it later costs trust.

Show the awkward result. A method that makes one project fall 13 places is more convincing when you point at it yourself.


Final scorecard

  • T1 + T2 claimed. 7/7 official checks PASS. T3 and T4 explicitly not claimed.
  • 273 tests pass, none skipped, standard library only.
  • Backend-enforced isolation, with 43 declared route policies checked against 5 identities.
  • Normalization: documented formula, k = 3, 32 of 41 projects re-ranked, Spearman 0.927 against raw.
  • Freshness gate (stale publish returns 409) and an immutable snapshot for explanations.
  • Seven audited defects, each with a regression test I watched fail.
  • Offline end to end: vendored wheels, seeded SQLite, no CDN, no client-side JS.
  • Known limitations written down, including the audit crash gap.

Reproduce it yourself

git clone https://github.com/rajpriyanid-creator/dogfood
cd dogfood
docker compose up              # portal at http://localhost:8080
python3 run.py .dogfood.toml   # the official checker, unmodified
python3 -m unittest discover -s tests -p "test_*.py"
Enter fullscreen mode Exit fullscreen mode
  • Repo: github.com/rajpriyanid-creator/dogfood
  • Read first: JUDGING.md (formula, worked fixture examples, limits), SECURITY.md (threat by threat, with residual limits), ARCHITECTURE.md (lifecycle, design decisions, how the tests were validated)
  • Event: dogfoodhack.com

Built for DOGFOOD 2026 by Rajpriyan S . Thanks to @hackathonraptors for a brief that rewards honesty over feature counts.

DOGFOOD #HackathonRaptors #Python #Security #OpenSource

Top comments (0)