DEV Community

Cover image for OUR DOGFOOD
ᴘᴀᴜʟ ғʀᴀɴᴄɪs
ᴘᴀᴜʟ ғʀᴀɴᴄɪs

Posted on

OUR DOGFOOD

We Thought Hackathon Judging Was Just About Scores. We Were Wrong.

A hackathon platform sounds simple on paper.

Create an event.

Accept submissions.

Assign judges.

Collect scores.

Rank teams.

Announce winners.

That was our mental model when we started building our platform for DOGFOOD 2026.

Then we started implementing the judging system.

Suddenly, the questions became much harder:

Who should judge which submission?

What happens when a judge is not eligible for a track?

How do we guarantee that every submission gets exactly the required number of reviews?

How do we prevent the same judge from reviewing the same submission twice?

How do we distribute work fairly between judges?

What does fair scoring even mean when one judge consistently gives 90s while another rarely goes above 70?

And what changes when an event has multiple tracks or multiple judging stages?

That was the point where our project stopped feeling like a conventional CRUD application.

We were no longer just building a hackathon platform.

We were building a judging system.

The platform we built around the judging workflow.


The Architecture Decision That Shaped Everything

The biggest decision we made was to treat judging as a pipeline, rather than a collection of separate features.

Our judging flow became:

Assignment
    ↓
Feasibility
    ↓
Workload Balancing
    ↓
Overlap Optimization
    ↓
Calibration
    ↓
Ranking
    ↓
Validation
Enter fullscreen mode Exit fullscreen mode

Every stage had a responsibility.

Assignment determines who evaluates what.

Feasibility checks whether that assignment is even possible.

Workload balancing prevents a few judges from receiving most of the work.

Overlap optimization makes sure submissions can be evaluated by multiple judges in a meaningful way.

Calibration attempts to compensate for systematic differences between judges.

Ranking converts the resulting scores into outcomes.

Validation checks the final state before anything is persisted.

That separation turned out to be extremely important.

It meant we could reason about each part independently instead of hiding the entire judging process inside one giant algorithm.


See It in Action

Here's a quick walkthrough of the platform, from event configuration and judge assignment to the judging workflow and results.

Judge Assignment Was More Complicated Than It Looked

Suppose there are:

  • N submissions
  • J eligible judges
  • R required reviews per submission

The first question is not:

"Which judge should get this submission?"

It is:

"Can the requested assignment even exist?"

We enforce constraints such as:

R ≤ J
Enter fullscreen mode Exit fullscreen mode

along with eligibility and duplicate-review checks.

Only after the assignment is feasible do we distribute the workload.

For a total of:

A = N × R
Enter fullscreen mode Exit fullscreen mode

reviews, the ideal workload is:

q = A / J
e = A mod J
Enter fullscreen mode Exit fullscreen mode

That means some judges receive q + 1 reviews and the remainder receive q.

It sounds simple.

It wasn't.

Because evenly distributing the number of reviews is only one part of the problem.

The assignments also need to respect:

  • Track eligibility
  • Judge availability
  • Required review count
  • Duplicate prevention
  • Intended overlap between judges

And that is why we eventually separated feasibility, assignment, and optimization instead of trying to solve everything in one step.


Two Judging Modes, Two Different Problems

One of the details we initially underestimated was the difference between track-based and non-track events.

They look similar from the UI.

Algorithmically, they are not.

Type 1 — No Tracks

All submissions belong to one global pool.

Judging can therefore operate across the entire submission set.

Type 2 — Multi-Track

Submissions are split into independent track pools.

Now judges can be eligible for some tracks and not others.

That means assignment has to happen inside each track, rather than treating the entire event as one giant pool.

This distinction affects everything downstream:

  • Assignment
  • Workload
  • Judge eligibility
  • Results
  • Overall ranking

That led us to explicitly model these two paths instead of forcing one generic algorithm to handle both.

Figure 1 — The judge assignment process is treated as a constrained pipeline rather than random allocation.


The Part We Were Surprisingly Proud Of: Overlap

At first, assigning multiple judges to a submission sounds trivial.

Just pick R judges.

But there is a hidden problem.

Suppose Judge A evaluates:

1, 2, 3, 4
Enter fullscreen mode Exit fullscreen mode

and Judge B evaluates:

5, 6, 7, 8
Enter fullscreen mode Exit fullscreen mode

There is no overlap between them.

Now imagine their scoring styles are very different.

How do we know whether that difference comes from the submissions or from the judges?

This is where intentional overlap becomes useful.

By making judges share some submissions, the system creates a connection between their scoring behaviour.

That overlap becomes valuable later when trying to normalize or calibrate scores.

This is one of those things that sounds obvious after you understand it.

We didn't understand its importance at the beginning.


We Didn't Want Raw Scores to Define the Ranking

This was probably the most interesting part of the system.

Imagine two judges.

Judge A tends to score aggressively:

82, 88, 91, 95
Enter fullscreen mode Exit fullscreen mode

Judge B tends to be much stricter:

58, 64, 69, 72
Enter fullscreen mode Exit fullscreen mode

A naive scoring system treats those values as directly comparable.

But the difference may not entirely represent project quality.

Some of it may simply represent judge behaviour.

So we implemented a Weighted Least Squares (WLS) calibration approach to account for systematic differences in judging.

Conceptually, the model considers something like:

Observed Score = Project Quality + Judge Effect + Noise
Enter fullscreen mode Exit fullscreen mode

The goal is not to tell judges how they should score.

It is to reduce the influence of systematic scoring differences when producing comparable results.

And this is exactly why overlap matters.

Without shared submissions between judges, there is much less information available to estimate how their scoring behaviour differs.

So assignment and calibration were not independent features anymore.

The assignment algorithm was helping create the data required by the calibration algorithm.

The algorithm preview exposing the judging pipeline and calibrated results.

That realization changed how we thought about the whole architecture.


The Judging System Became a Pipeline

One thing we deliberately avoided was treating judging, normalization, and ranking as completely separate features.

Instead, we designed them as connected stages.

A review produces data.

Assignments determine who produced that data.

Overlap creates relationships between judges.

Calibration transforms the scores.

Ranking consumes the calibrated results.

And validation makes sure the final state still satisfies the rules.

So the system becomes:

Submissions
     ↓
Judge Assignment
     ↓
Reviews
     ↓
Calibration / Normalisation
     ↓
Ranking
     ↓
Results
Enter fullscreen mode Exit fullscreen mode

This sounds straightforward when written as a diagram.

Implementing it without breaking one stage with another was much less straightforward.


Multi-Stage Judging Made the Problem Bigger

Another assumption we had was that an event would have one judging round.

Then we thought about actual hackathons.

Some events may have:

  • An initial screening round
  • A technical evaluation
  • A final presentation
  • A grand judging round

So we designed the architecture around multiple judging stages.

That means the system cannot simply store:

"Project X has been judged."

It needs to understand:

"Project X was judged in Stage 1 using this configuration, then evaluated again in Stage 2 using another judging setup."

That sounds like a small schema change.

It isn't.

Once stages exist, assignment, rubrics, results, ranking and progression all need to understand where in the judging pipeline a submission currently belongs.

That was one of the moments where we realized that seemingly small product decisions can have major architectural consequences.

Figure 2 — Track-based and non-track events follow different judging paths before producing final results.


What We Did Well

1. We Separated Responsibilities Early

Instead of building one giant judging function, we broke the problem into smaller conceptual components:

Feasibility

Can the assignment be made?

Deterministic Assignment

Can we produce exactly R reviews per submission while respecting eligibility and preventing duplicates?

Workload Balancing

Can the reviews be distributed fairly?

Overlap Optimization

Can we improve the connectivity between judges?

Calibration

Can we account for systematic differences in judging?

Validation

Did the final result still satisfy every constraint?

That decomposition made an otherwise complicated problem much easier to reason about.


2. We Treated Permissions as Part of the Domain

The platform has clearly separated roles:

Admin → Organizer → Judge → Participant

Each role has its own responsibilities and boundaries.

We didn't want authorization to be something that existed only in the frontend.

A judge shouldn't simply not see something.

The system should actually prevent them from accessing things they are not supposed to access.

That distinction became increasingly important as judging logic became more complex.

Role-aware platform administration and user management.


3. We Treated Track and Non-Track Judging as Different Algorithmic Problems

Instead of forcing every event through the same ranking and assignment flow, we explicitly considered:

Global submission pool for non-track events.

Independent track pools for multi-track events.

That made judge eligibility and assignment much more predictable.


4. We Connected Assignment With Calibration

This was probably our most important architectural insight.

At first glance, these look like separate problems:

Assign judges.

Normalize scores.

In reality, they are connected.

The assignment strategy determines which judges overlap on which submissions.

That overlap determines how much information exists for calibration.

So the first algorithm influences the quality of the second.


Then Docker Happened.

Not the interesting kind of happened.

The painful kind.

Our application had to run reliably through Docker, including under the constraints of the competition environment.

The initial Docker setup itself wasn't particularly complicated.

Then we discovered why our builds were taking absurdly long.

20–30 Minutes.

For a build.

We started looking at the application.

Then the dependencies.

Then Docker itself.

Then everything else.

The actual problem was embarrassingly simple.

We had a large testing ZIP file sitting inside the project, and we had forgotten to exclude it through .dockerignore.

Docker was sending that file as part of the build context.

So every build was dragging around a file that had absolutely no reason to be there.

One missing .dockerignore entry turned a normal development cycle into:

Make change
    ↓
docker compose up
    ↓
Wait
    ↓
Wait
    ↓
Wait
    ↓
Wonder what is happening
Enter fullscreen mode Exit fullscreen mode

Once we found it, the fix was tiny.

The debugging time wasn't.


The Docker Lesson Wasn't "Use .dockerignore"

That would be too easy.

The actual lesson was:

Build context is part of your build system.

When a Docker build is slow, it is tempting to immediately inspect dependencies, image layers, package installation, or application startup.

Sometimes the problem happens before any of those things.

The files you send to Docker matter.

A massive build context can make the rest of your optimization efforts irrelevant.

We learned that the hard way.


Another Weird Problem: Authentication Tests

We also ran into inconsistent authentication test behaviour.

The same run.py tests could behave differently across systems.

One system would pass authentication-related tests.

Another would fail them.

That created a particularly annoying debugging problem:

Is the application broken?

or

Is the environment different?

This was a useful reminder that reproducibility is not just about having the same source code.

Environment state matters.

Dependencies matter.

Configuration matters.

And when you're debugging during a hackathon, distinguishing between an application bug and an environment problem can consume an unreasonable amount of time.


What We Didn't Finish

Not everything made it into the final implementation.

Our proposed judging architecture included pairwise judging with the Bradley–Terry model.

We liked the idea because pairwise comparisons ask a different question.

Instead of asking:

"How many points does this project deserve?"

they ask:

"Which of these two projects is better?"

From those comparisons, a ranking model can estimate relative strength.

We had designed this as a potential judging stage.

But we did not fully implement it as an organizer-selectable option within the available time.

And we're actually glad we didn't rush it.

A half-working statistical model buried inside the judging pipeline would have been much worse than explicitly leaving it unfinished.


The UI Lost the Race Against the Algorithm

There was another tradeoff we had to make.

We wanted a more polished UI.

We didn't have unlimited time.

Eventually we had to decide:

Do we spend another few hours polishing visual details, or make sure the judging engine actually behaves correctly?

We chose the latter.

That meant sacrificing some visual polish in favour of getting the core judging workflow right.

The interface is usable.

It isn't the most visually sophisticated part of the project.

But the system underneath it is much more interesting than the pixels suggest.

And for this project, that felt like the right tradeoff.


The Biggest Thing We Learned

Before building this, we thought a hackathon platform was mainly a collection of features.

Authentication.

Submissions.

Judges.

Scores.

Leaderboard.

After building the judging system, we think about it differently.

The difficult part isn't storing a score.

It's deciding how that score should be produced, who should produce it, what constraints apply, how different evaluators affect it, and how the final result should be constructed.

The leaderboard is just the final visible number.

Underneath it is a chain of decisions.

And every decision can introduce bias, inconsistency, or failure.

That changed the way we approached the entire system.


What We Would Do Differently Next Time

We would spend more time on the infrastructure around the algorithm before starting implementation.

We would make the Docker build context explicit from day one.

We would test authentication in more controlled environments earlier.

We would build stronger test cases around judge assignment edge cases.

And most importantly, we would formalize the judging model before writing too much application code.

Because the hardest bugs weren't UI bugs.

They weren't even ordinary backend bugs.

They were modeling bugs.

We were trying to turn a real-world judging process—with constraints, people, preferences, fairness, and incomplete information—into deterministic software.

That is a much harder problem than it first appears.


The Architecture We Ended Up With

At a high level, our judging pipeline became:

Event
  ↓
Judging Configuration
  ↓
Track / Global Submission Pool
  ↓
Eligible Judges
  ↓
Feasibility Checks
  ↓
Deterministic Assignment
  ↓
Workload Balancing
  ↓
Overlap Optimization
  ↓
Judging
  ↓
WLS Calibration
  ↓
Ranking
  ↓
Validation
  ↓
Persisted Results
Enter fullscreen mode Exit fullscreen mode

The important part isn't that the pipeline looks neat.

It's that each stage exists because we encountered a reason it needed to exist.


Final Takeaway

We started with:

"Let's build a hackathon platform."

We ended up spending a surprising amount of time thinking about:

allocation, fairness, statistical calibration, constraints, reproducibility, and failure modes.

And that's probably the biggest lesson we got from DOGFOOD 2026.

Building the screens is the visible part of a hackathon platform.

The difficult part is deciding what should happen after everyone clicks Submit.

Because once the submissions are in, the system has to make decisions.

And those decisions need to be:

deterministic enough to trust, flexible enough to configure, and explainable enough to defend.

That was the real system we were trying to build.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more