Hackathon Raptors gave every team at DOGFOOD 2026 the same brief and 72 hours: build the submission-and-judging platform they would actually run. The winner gets forked into production.
We're a team of two. Our submission is Dogfood — 9 Spring Boot microservices behind a gateway, a React frontend, PostgreSQL with Row-Level Security, RabbitMQ for the async event mesh, Redis, MinIO, and a judging engine built around z-score normalization with Bayesian shrinkage.
This is the question we kept returning to while building it:
Where is this rule actually enforced?
Not "does the UI hide the button." Not "does the frontend route guard redirect you." Where in the stack does the rule live, such that bypassing it is architecturally impossible?
That question shaped every significant decision we made — including which tools we reached for and why.
Hiding a rule is not the same as enforcing it
The hackathon spec says it plainly: role isolation must be enforced in the backend, and if it only works because the UI hides a button, it does not work.
The naive version of "judges can't see each other's scores" is a frontend route guard plus a controller that trusts a judge_id from the request body. Anyone with curl and a valid token can swap the ID and get another judge's scores back.
We enforce it twice, and the second layer is the one that matters:
Layer 1 — Application: every score query filters by judge_id taken from the verified JWT claim (X-User-Id injected by the gateway), never from a request parameter.
Layer 2 — Database engine: PostgreSQL Row-Level Security keyed on a session variable set from that same claim before every query:
SET LOCAL app.current_user_id = '<uuid-from-jwt>';
CREATE POLICY judge_own_scores ON judging.scores
USING (
judge_id = current_setting('app.current_user_id')::uuid
OR current_setting('app.current_role', true) IN ('ORGANIZER', 'ADMIN')
);
Even if the application code has a bug, the PostgreSQL kernel silently drops rows that don't belong to the requesting judge. The acceptance suite's T2 check — judge cannot see peer scores — passes not because we wrote a special case for it, but because the rule lives at a layer that the request can't talk its way around.
The deadline is a domain constraint, not a UI feature
A deadline shown in the frontend is a suggestion. A deadline checked against a client-sent timestamp is a suggestion with a race condition.
The Submission Service enforces the cutoff on the server with a Redis distributed lock (SETNX) and compares against the event's stored deadline using server time. The client's clock is never trusted. A submission attempted after the deadline is rejected regardless of client clock skew or network latency.
The rule lives where authoritative time lives.
The math problem hiding inside score averaging
Consider three judges scoring the same projects:
- Judge A (Generous): scores 8, 9, 10. Mean = 9, StdDev = 1.
- Judge B (Harsh): scores 2, 3, 4. Mean = 3, StdDev = 1.
- Judge C (Neutral): scores 4, 6, 8. Mean = 6, StdDev = 2.
Project X gets a 9 from Judge A, a 4 from Judge B, and an 8 from Judge C. Raw average: 7.0.
But look at what those scores actually mean relative to each judge's baseline:
- Judge A's z-score for X:
(9 - 9) / 1 = 0— exactly average for them - Judge B's z-score for X:
(4 - 3) / 1 = +1— above average for them - Judge C's z-score for X:
(8 - 6) / 2 = +1— above average for them
Normalized average: +0.67. Judge A gave the highest raw score but contributed zero ranking signal — it was just their baseline. Judge B gave a low raw score but it was actually a positive signal relative to how they score everything.
This is what per-judge, per-criterion z-score normalization does:
z = (raw_score − judge_mean) / judge_std_dev
It corrects both a judge's central tendency (hawk vs. dove) and their spread (a judge who clusters scores in a narrow band). Min-max scaling only fixes the range — z-score corrects central tendency too.
Bayesian shrinkage: what to do when a judge has only reviewed 2 projects
Z-scores have a failure mode: a judge with 1 or 2 reviews gives you noisy estimates of their own mean and standard deviation. If that judge happens to draw a genuinely strong project, their normalized score can swing the leaderboard disproportionately.
We apply Bayesian shrinkage with k₀ = 5 to handle this:
λ = k / (k + 5)
z_shrunk = λ × z + (1 − λ) × global_mean
- Judge with 1 review: λ = 1/6 ≈ 0.167 — their score is dampened 83.3% toward the global consensus mean.
- Judge with 15 reviews: λ = 15/20 = 0.75 — only 25% dampening, their calibration is trusted.
One edge case handled explicitly: a judge who gives every project the same score has σ = 0, causing division by zero in the z-score formula. We set σ = 1.0 in that case, which maps all their scores to z = 0 and contributes zero ranking signal. This is a verified invariant in NormalizationEngineTest.java.
The hawk/dove proof falls out cleanly: if Judge A (mean 8, σ 1) and Judge B (mean 4, σ 1) both score a standout project at +2σ above their own baseline (Judge A gives 10, Judge B gives 6), both produce z = +2.0. Their leniency difference becomes mathematically irrelevant.
Catching pre-built projects without shelling out
One integrity problem most platforms we studied don't address: a team builds most of the project before the event, pushes to a private repo, then makes one commit during it. The submission looks legitimate.
The Submission Service runs an in-memory Eclipse JGit scan on every submitted repository URL and records a risk score:
0 commits → risk 100 (EMPTY_REPOSITORY)
1 commit → risk 80 (SINGLE_COMMIT_DUMP)
>90% commits before event start → HIGH (PRE_HACKATHON_CODE)
Two decisions worth explaining:
Pure Java, no shell commands. Running git clone with a user-supplied URL is a command-injection surface. Eclipse JGit in a Java memory buffer removes it.
A signal, not a verdict. The scanner puts anomalies in an organizer review queue and never auto-disqualifies. A one-commit repo might be a legitimate solo build with poor git habits. Detection belongs to the machine; the decision belongs to a human.
Why not a monolith? Why RabbitMQ? Why Redis?
Fair questions for a 72-hour hackathon, where the simpler choice is usually the right one.
On microservices vs. monolith: the decision wasn't about scale. At fixture size — 40 projects, 30 judges — a well-structured monolith would have been fine. We split into services because the security boundaries are real and need to be structural, not just logical. The Submission Service must not be able to touch the judging schema. The Judging Service must not be able to write to submissions. When those concerns share a codebase and a connection pool, you're trusting developer discipline to maintain the boundary. When they're separate services with separate Flyway-managed schemas and no shared connection, the boundary is physical. You can't accidentally cross it.
On RabbitMQ: score submission is a write path that needs to stay fast, especially in the final hour of an event when judges are racing the deadline. If submitting a score also triggered synchronous re-normalization across all scores for the event, that write would slow down as the event grew. Instead, the score commits to PostgreSQL and publishes a score.submitted event to the dogfood.events topic exchange in under 50ms. Normalization recompute, audit log write, webhook dispatch, and notification all happen asynchronously on their own queues — audit.queue, certificate.queue, webhook.queue, notification.queue. The judge gets a fast response regardless of what follows.
On Redis: two separate jobs. The Voting Service rate-limiter uses a Redis token-bucket per IP/account — Redis's atomic increment operations make this correct under concurrency in a way a database counter isn't, without the lock overhead. The submission deadline lock uses SETNX because it needs to be atomic and shared across any instance of the Submission Service. Both are cases where Redis does something PostgreSQL can technically do, but with better latency and simpler concurrency guarantees for their specific access patterns.
What the architecture diagram doesn't show
Our README has an ASCII diagram of all 9 services. What it doesn't show is that the diagram is the output of one question asked repeatedly: where does this rule live, and can someone get around that layer from outside?
- Deadline enforcement → Redis lock on the server, not a client timestamp
- Score isolation → PostgreSQL RLS at the kernel, not a route guard
- Score fairness → Bayesian normalization in the judging engine, not a post-event spreadsheet
- Commit integrity → in-memory JGit at submission, not an organizer eyeballing GitHub profiles
- Webhook trust → HMAC-SHA256 signed payloads, not plain HTTP callbacks
- Concurrency safety → RabbitMQ decoupling, not synchronous blocking writes
The microservices weren't the goal. They fell out of keeping each rule in the layer that owns it.
The moment the architecture pushed back
One thing no diagram captures: docker compose up ran out of resources on a local machine.
By the time all containers were up, the stack was consuming 3.13GB of RAM on a 7.4GB host and peaking at 1376% CPU across 12 cores — all before a single HTTP request had landed. Nine Spring Boot services each with their own JVM heap, plus PostgreSQL, RabbitMQ, Redis, and MinIO, adds up fast.
We moved to a cloud shell instance, which was only possible because my teammate had a student subscription. Without it, we would have been cutting services under time pressure.
That's an honest cost of the microservices decision, and we'd rather say it than hide it. The fix — capping each service's JVM heap with -Xmx in the Compose file — is straightforward. Most of these services don't need 512MB each at hackathon scale. But we ran out of time before we could tune it, and the spec asks for something that runs on a laptop. Ours runs, but comfortably only on a machine with enough headroom, and that's worth saying plainly.
What we claimed and what we proved
.dogfood.toml claims T1 and T2. The acceptance suite verifies exactly that — 7 of 7 checks passing:
T1 gallery is public ................. PASS
T1 project from fixtures shown ....... PASS
T1 closed event refuses submissions .. PASS
T2 judge sees own scores ............. PASS
T2 judge cannot see peer scores ...... PASS
T2 participant blocked ............... PASS
T2 csv export works .................. PASS
claimed T1 T2, verified T1 T2
Quadratic voting, Ed25519-signed certificates, the OpenAPI spec, and bulk export are implemented, but the acceptance suite doesn't cover them, so we don't claim them in .dogfood.toml.
A judging platform that is vague about what it can prove is hard to trust. We tried to be exact about our own limits.
Repo: github.com/codewisp-ai/DogFood
Stack: Java 21 · Spring Boot 3.3 · PostgreSQL 16 · RabbitMQ 3.13 · Redis 7 · MinIO · React
License: MIT · docker compose up -d
Built for DOGFOOD 2026 by Hackathon Raptors. #Dogfood2026
Top comments (0)