When we started building TrustForge for DogFood 2026, I assumed the judging dashboard would be the hard part. It turned out to be the easy part. What actually took thinking was a question I hadn't planned for: how do you explain a hackathon result after it's already been published?
Most judging platforms store scores and pick a winner. We wanted the process behind the winner to be inspectable too, and that one decision ended up shaping almost everything else.
Why we stayed with a monolith
My first instinct was to split it all up: auth, judging, assignment, scoring, audit and notifications, each as its own service. It sounds scalable, but for a hackathon it would have been a lot of overhead for no real gain.
So TrustForge is one Spring Boot app with clearly separated modules (authentication, authorization, submissions, judging, normalization, anomalies, audit, results), and a separate React frontend that talks to versioned REST APIs. I cared less about "monolith vs microservices" than about keeping the boundaries explicit without taking on distributed-systems problems we didn't have.
Everything runs through one pipeline, and it became the mental model for the whole project:
SubmissionVersion → Assignment → Evaluation → NormalizationRun
→ Anomaly → AuditEvent → ResultSnapshot
"Assign judges to projects" is not a simple requirement
On paper it's one line. In practice you have to handle judge capacity, minimum coverage per project, declared conflicts, workload balance, and the fact that rerunning the algorithm should give the same answer.
We also didn't want an assignment to just say Judge A → Project 12. We wanted it to carry its reasoning:
Judge A → Project 12
Eligible: yes | Conflict: none | Capacity: 12/20
Coverage: satisfied | Fairness: 0.94
Algorithm: v1.2 | Seed: 8f91c7
The seed and algorithm version matter more than they look. If an organizer reruns the assignment and gets something completely different, you can't debug it and you definitely can't explain it. Every run records both, and the acceptance tests check that declared conflicts are excluded and coverage is met.
Judges don't use the same scale
Say Judge A usually gives 8.5, 9, 9.5, 10 and Judge B usually gives 6, 6.5, 7, 7.5. If you just average, Judge A's habits carry more weight, even though neither judge is wrong. The numbers look comparable, but they aren't.
So we added per-judge normalization:
z = (score - judge mean) / judge standard deviation
We kept the raw score. Normalization is an interpretation of the data, and I didn't want to throw away the original evidence just because we'd built a transformed version. A judge with zero standard deviation is handled explicitly instead of crashing into a divide-by-zero, and missing evaluations are left missing, never filled in.
The scoring formula was wrong, and it looked fine
This is probably the most useful thing I learned on the project.
Our methodology says the result combines normalized judging with community voting. While reviewing the implementation, we noticed the two inputs weren't on compatible scales. One was a z-score and the other was a raw vote count. Adding them directly gives you something that looks sophisticated and quietly lets one side dominate.
That made us ask what "80% judging, 20% community" actually means. It only means something if both signals are mapped to the same range first:
final = (judging, scaled 0–100) × 0.80
+ (community, scaled 0–100) × 0.20
The takeaway went beyond this one formula: a score can look reasonable in the UI and still be mathematically off, so check the units, the ranges, the edge cases, and whether the code matches the documented method. It's also why the results page shows the methodology instead of just a final number.
An audit table isn't the same as a trustworthy audit log
A normal audit table tells you what's stored. It doesn't tell you whether someone edited the history.
So we chained the audit events with SHA-256, starting from a GENESIS value, where each event's hash includes the previous one. Each record stores the previous hash, current hash, actor, action, entity, timestamp, request ID and payload. Verification recomputes the whole chain from the start.
We also wrote a test that deliberately edits an earlier payload and confirms verification fails. I trust that test far more than one that only checks the happy path returns true.
Security lives on the backend
It's easy to treat a hidden button as access control. if (user.role === "ORGANIZER") showButton() is a UX choice, not a security boundary. The backend decides what's allowed:
- Organizer: assignment management
- Judge: their own assigned evaluations
- Participant: the public gallery and voting
In acceptance testing, a judge trying to reach organizer-only assignments got a 403. Access tokens expire, refresh tokens rotate, and reusing an old refresh token after rotation gets rejected.
Docker, and the difference between "written" and "run"
Docker didn't magically solve deployment. The more useful lesson was keeping the design of the deployment separate from the verification of it.
The repo has a Compose setup with PostgreSQL, Redis, the Spring Boot API, and Nginx serving the React frontend, with health checks gating the startup order. But our acceptance environment didn't have Docker, so we didn't call Compose verified. The acceptance report has two buckets: VERIFIED and NOT VERIFIED / BLOCKED BY ENVIRONMENT.
It's a small documentation choice, but it's one of the easiest ways to stop a project looking more finished than it is. "I wrote the config" and "I ran the config" are not the same sentence.
The acceptance suite became part of the product
We didn't want the demo to rest on "trust me, it works." The smoke test covers the seeded gallery, login, dashboard, assignment coverage, normalization, audit verification, results, certificate verification, and role isolation. The live run passed ten checks against the local API. We added focused tests for deterministic normalization, audit tamper detection, assignment conflicts and capacity, and duplicate voting.
Somewhere along the way the suite stopped being just a test suite and became another way of stating what TrustForge promises.
What I'd change
A runnable demo can make the underlying model look finished, and ours isn't. The demo uses a deterministic in-memory store as a replaceable persistence layer. PostgreSQL and Flyway are part of the deployment design, but the demo shouldn't be mistaken for a production persistence setup.
If I built it again, I'd:
- Move the judging model fully into PostgreSQL
- Persist assignment runs and normalization datasets
- Make result snapshots immutable at the database level
- Build a real pairwise ranking model in place of the current read-model placeholder
- Add property-based tests around assignment and scoring
- Make the final weights and normalization ranges explicit in code and tests
None of these are hidden gaps. They're documented.
What it changed for me
Before this, my question for a feature was "does it work?" TrustForge added two more: can someone understand why it worked, and can they verify that explanation without taking my word for it?
Those questions changed the architecture, the tests and the UI. They also changed how I think about trust in software: you can't bolt it on at the end. Reproducibility, explainability, versioning, testability and auditability have to be there from the first decision.
So the idea behind TrustForge is simple: don't just publish the result, preserve the process that produced it.
Top comments (0)