<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ᴘᴀᴜʟ ғʀᴀɴᴄɪs</title>
    <description>The latest articles on DEV Community by ᴘᴀᴜʟ ғʀᴀɴᴄɪs (@_s_619774fa3fd2).</description>
    <link>https://dev.to/_s_619774fa3fd2</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4156068%2Fd25d8f00-b346-4c38-a01a-084d1ef1826a.jpg</url>
      <title>DEV Community: ᴘᴀᴜʟ ғʀᴀɴᴄɪs</title>
      <link>https://dev.to/_s_619774fa3fd2</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/_s_619774fa3fd2"/>
    <language>en</language>
    <item>
      <title>OUR DOGFOOD</title>
      <dc:creator>ᴘᴀᴜʟ ғʀᴀɴᴄɪs</dc:creator>
      <pubDate>Thu, 08 Oct 2026 17:41:43 +0000</pubDate>
      <link>https://dev.to/_s_619774fa3fd2/our-dogfood-2a90</link>
      <guid>https://dev.to/_s_619774fa3fd2/our-dogfood-2a90</guid>
      <description>&lt;h2&gt;
  
  
  We Thought Hackathon Judging Was Just About Scores. We Were Wrong.
&lt;/h2&gt;

&lt;p&gt;A hackathon platform sounds simple on paper.&lt;/p&gt;

&lt;p&gt;Create an event.&lt;br&gt;&lt;br&gt;
Accept submissions.&lt;br&gt;&lt;br&gt;
Assign judges.&lt;br&gt;&lt;br&gt;
Collect scores.&lt;br&gt;&lt;br&gt;
Rank teams.&lt;br&gt;&lt;br&gt;
Announce winners.&lt;/p&gt;

&lt;p&gt;That was our mental model when we started building our platform for &lt;strong&gt;DOGFOOD 2026&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Then we started implementing the judging system.&lt;/p&gt;

&lt;p&gt;Suddenly, the questions became much harder:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Who should judge which submission?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What happens when a judge is not eligible for a track?&lt;/p&gt;

&lt;p&gt;How do we guarantee that every submission gets exactly the required number of reviews?&lt;/p&gt;

&lt;p&gt;How do we prevent the same judge from reviewing the same submission twice?&lt;/p&gt;

&lt;p&gt;How do we distribute work fairly between judges?&lt;/p&gt;

&lt;p&gt;What does &lt;strong&gt;fair scoring&lt;/strong&gt; even mean when one judge consistently gives 90s while another rarely goes above 70?&lt;/p&gt;

&lt;p&gt;And what changes when an event has multiple tracks or multiple judging stages?&lt;/p&gt;

&lt;p&gt;That was the point where our project stopped feeling like a conventional CRUD application.&lt;/p&gt;

&lt;p&gt;We were no longer just building a hackathon platform.&lt;/p&gt;

&lt;p&gt;We were building a &lt;strong&gt;judging system&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbv30kwt4tzv8g6ykwhl3.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbv30kwt4tzv8g6ykwhl3.jpeg" alt="The platform we built around the judging workflow." width="800" height="380"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  The Architecture Decision That Shaped Everything
&lt;/h2&gt;

&lt;p&gt;The biggest decision we made was to treat judging as a &lt;strong&gt;pipeline&lt;/strong&gt;, rather than a collection of separate features.&lt;/p&gt;

&lt;p&gt;Our judging flow became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Assignment
    ↓
Feasibility
    ↓
Workload Balancing
    ↓
Overlap Optimization
    ↓
Calibration
    ↓
Ranking
    ↓
Validation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every stage had a responsibility.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Assignment&lt;/strong&gt; determines &lt;em&gt;who evaluates what&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feasibility&lt;/strong&gt; checks whether that assignment is even possible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Workload balancing&lt;/strong&gt; prevents a few judges from receiving most of the work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Overlap optimization&lt;/strong&gt; makes sure submissions can be evaluated by multiple judges in a meaningful way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Calibration&lt;/strong&gt; attempts to compensate for systematic differences between judges.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ranking&lt;/strong&gt; converts the resulting scores into outcomes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Validation&lt;/strong&gt; checks the final state before anything is persisted.&lt;/p&gt;

&lt;p&gt;That separation turned out to be extremely important.&lt;/p&gt;

&lt;p&gt;It meant we could reason about each part independently instead of hiding the entire judging process inside one giant algorithm.&lt;/p&gt;




&lt;h2&gt;
  
  
  See It in Action
&lt;/h2&gt;

&lt;p&gt;Here's a quick walkthrough of the platform, from event configuration and judge assignment to the judging workflow and results.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/0lQymqb-C1M" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  Judge Assignment Was More Complicated Than It Looked
&lt;/h2&gt;

&lt;p&gt;Suppose there are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;N&lt;/strong&gt; submissions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;J&lt;/strong&gt; eligible judges&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;R&lt;/strong&gt; required reviews per submission&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first question is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Which judge should get this submission?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Can the requested assignment even exist?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;We enforce constraints such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;R ≤ J
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;along with eligibility and duplicate-review checks.&lt;/p&gt;

&lt;p&gt;Only after the assignment is feasible do we distribute the workload.&lt;/p&gt;

&lt;p&gt;For a total of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A = N × R
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;reviews, the ideal workload is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;q = A / J
e = A mod J
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That means some judges receive &lt;code&gt;q + 1&lt;/code&gt; reviews and the remainder receive &lt;code&gt;q&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It sounds simple.&lt;/p&gt;

&lt;p&gt;It wasn't.&lt;/p&gt;

&lt;p&gt;Because evenly distributing the number of reviews is only one part of the problem.&lt;/p&gt;

&lt;p&gt;The assignments also need to respect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Track eligibility&lt;/li&gt;
&lt;li&gt;Judge availability&lt;/li&gt;
&lt;li&gt;Required review count&lt;/li&gt;
&lt;li&gt;Duplicate prevention&lt;/li&gt;
&lt;li&gt;Intended overlap between judges&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And that is why we eventually separated &lt;strong&gt;feasibility&lt;/strong&gt;, &lt;strong&gt;assignment&lt;/strong&gt;, and &lt;strong&gt;optimization&lt;/strong&gt; instead of trying to solve everything in one step.&lt;/p&gt;




&lt;h2&gt;
  
  
  Two Judging Modes, Two Different Problems
&lt;/h2&gt;

&lt;p&gt;One of the details we initially underestimated was the difference between &lt;strong&gt;track-based&lt;/strong&gt; and &lt;strong&gt;non-track&lt;/strong&gt; events.&lt;/p&gt;

&lt;p&gt;They look similar from the UI.&lt;/p&gt;

&lt;p&gt;Algorithmically, they are not.&lt;/p&gt;

&lt;h3&gt;
  
  
  Type 1 — No Tracks
&lt;/h3&gt;

&lt;p&gt;All submissions belong to one global pool.&lt;/p&gt;

&lt;p&gt;Judging can therefore operate across the entire submission set.&lt;/p&gt;

&lt;h3&gt;
  
  
  Type 2 — Multi-Track
&lt;/h3&gt;

&lt;p&gt;Submissions are split into independent track pools.&lt;/p&gt;

&lt;p&gt;Now judges can be eligible for some tracks and not others.&lt;/p&gt;

&lt;p&gt;That means assignment has to happen &lt;strong&gt;inside each track&lt;/strong&gt;, rather than treating the entire event as one giant pool.&lt;/p&gt;

&lt;p&gt;This distinction affects everything downstream:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Assignment&lt;/li&gt;
&lt;li&gt;Workload&lt;/li&gt;
&lt;li&gt;Judge eligibility&lt;/li&gt;
&lt;li&gt;Results&lt;/li&gt;
&lt;li&gt;Overall ranking&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That led us to explicitly model these two paths instead of forcing one generic algorithm to handle both.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftkbvrl203mz7fonf6uv0.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftkbvrl203mz7fonf6uv0.jpeg" width="770" height="1545"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1 — The judge assignment process is treated as a constrained pipeline rather than random allocation.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Part We Were Surprisingly Proud Of: Overlap
&lt;/h2&gt;

&lt;p&gt;At first, assigning multiple judges to a submission sounds trivial.&lt;/p&gt;

&lt;p&gt;Just pick &lt;code&gt;R&lt;/code&gt; judges.&lt;/p&gt;

&lt;p&gt;But there is a hidden problem.&lt;/p&gt;

&lt;p&gt;Suppose Judge A evaluates:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1, 2, 3, 4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and Judge B evaluates:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;5, 6, 7, 8
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no overlap between them.&lt;/p&gt;

&lt;p&gt;Now imagine their scoring styles are very different.&lt;/p&gt;

&lt;p&gt;How do we know whether that difference comes from the submissions or from the judges?&lt;/p&gt;

&lt;p&gt;This is where &lt;strong&gt;intentional overlap&lt;/strong&gt; becomes useful.&lt;/p&gt;

&lt;p&gt;By making judges share some submissions, the system creates a connection between their scoring behaviour.&lt;/p&gt;

&lt;p&gt;That overlap becomes valuable later when trying to normalize or calibrate scores.&lt;/p&gt;

&lt;p&gt;This is one of those things that sounds obvious after you understand it.&lt;/p&gt;

&lt;p&gt;We didn't understand its importance at the beginning.&lt;/p&gt;




&lt;h2&gt;
  
  
  We Didn't Want Raw Scores to Define the Ranking
&lt;/h2&gt;

&lt;p&gt;This was probably the most interesting part of the system.&lt;/p&gt;

&lt;p&gt;Imagine two judges.&lt;/p&gt;

&lt;p&gt;Judge A tends to score aggressively:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;82, 88, 91, 95
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Judge B tends to be much stricter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;58, 64, 69, 72
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A naive scoring system treats those values as directly comparable.&lt;/p&gt;

&lt;p&gt;But the difference may not entirely represent project quality.&lt;/p&gt;

&lt;p&gt;Some of it may simply represent &lt;strong&gt;judge behaviour&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So we implemented a &lt;strong&gt;Weighted Least Squares (WLS) calibration&lt;/strong&gt; approach to account for systematic differences in judging.&lt;/p&gt;

&lt;p&gt;Conceptually, the model considers something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Observed Score = Project Quality + Judge Effect + Noise
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not to tell judges how they should score.&lt;/p&gt;

&lt;p&gt;It is to reduce the influence of systematic scoring differences when producing comparable results.&lt;/p&gt;

&lt;p&gt;And this is exactly why overlap matters.&lt;/p&gt;

&lt;p&gt;Without shared submissions between judges, there is much less information available to estimate how their scoring behaviour differs.&lt;/p&gt;

&lt;p&gt;So assignment and calibration were not independent features anymore.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The assignment algorithm was helping create the data required by the calibration algorithm.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F63qo64ocx9nonwazitpk.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F63qo64ocx9nonwazitpk.jpeg" alt="The algorithm preview exposing the judging pipeline and calibrated results." width="800" height="371"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That realization changed how we thought about the whole architecture.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Judging System Became a Pipeline
&lt;/h2&gt;

&lt;p&gt;One thing we deliberately avoided was treating judging, normalization, and ranking as completely separate features.&lt;/p&gt;

&lt;p&gt;Instead, we designed them as connected stages.&lt;/p&gt;

&lt;p&gt;A review produces data.&lt;/p&gt;

&lt;p&gt;Assignments determine who produced that data.&lt;/p&gt;

&lt;p&gt;Overlap creates relationships between judges.&lt;/p&gt;

&lt;p&gt;Calibration transforms the scores.&lt;/p&gt;

&lt;p&gt;Ranking consumes the calibrated results.&lt;/p&gt;

&lt;p&gt;And validation makes sure the final state still satisfies the rules.&lt;/p&gt;

&lt;p&gt;So the system becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Submissions
     ↓
Judge Assignment
     ↓
Reviews
     ↓
Calibration / Normalisation
     ↓
Ranking
     ↓
Results
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This sounds straightforward when written as a diagram.&lt;/p&gt;

&lt;p&gt;Implementing it without breaking one stage with another was much less straightforward.&lt;/p&gt;




&lt;h2&gt;
  
  
  Multi-Stage Judging Made the Problem Bigger
&lt;/h2&gt;

&lt;p&gt;Another assumption we had was that an event would have one judging round.&lt;/p&gt;

&lt;p&gt;Then we thought about actual hackathons.&lt;/p&gt;

&lt;p&gt;Some events may have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An initial screening round&lt;/li&gt;
&lt;li&gt;A technical evaluation&lt;/li&gt;
&lt;li&gt;A final presentation&lt;/li&gt;
&lt;li&gt;A grand judging round&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So we designed the architecture around &lt;strong&gt;multiple judging stages&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That means the system cannot simply store:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Project X has been judged."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It needs to understand:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Project X was judged in Stage 1 using this configuration, then evaluated again in Stage 2 using another judging setup."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sounds like a small schema change.&lt;/p&gt;

&lt;p&gt;It isn't.&lt;/p&gt;

&lt;p&gt;Once stages exist, assignment, rubrics, results, ranking and progression all need to understand where in the judging pipeline a submission currently belongs.&lt;/p&gt;

&lt;p&gt;That was one of the moments where we realized that seemingly small product decisions can have major architectural consequences.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff7mwcr41alwt9bxthzmj.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff7mwcr41alwt9bxthzmj.jpeg" width="800" height="775"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 2 — Track-based and non-track events follow different judging paths before producing final results.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What We Did Well
&lt;/h2&gt;

&lt;h2&gt;
  
  
  1. We Separated Responsibilities Early
&lt;/h2&gt;

&lt;p&gt;Instead of building one giant judging function, we broke the problem into smaller conceptual components:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feasibility&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Can the assignment be made?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deterministic Assignment&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Can we produce exactly &lt;code&gt;R&lt;/code&gt; reviews per submission while respecting eligibility and preventing duplicates?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Workload Balancing&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Can the reviews be distributed fairly?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Overlap Optimization&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Can we improve the connectivity between judges?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Calibration&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Can we account for systematic differences in judging?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Validation&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Did the final result still satisfy every constraint?&lt;/p&gt;

&lt;p&gt;That decomposition made an otherwise complicated problem much easier to reason about.&lt;/p&gt;


&lt;h2&gt;
  
  
  2. We Treated Permissions as Part of the Domain
&lt;/h2&gt;

&lt;p&gt;The platform has clearly separated roles:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Admin → Organizer → Judge → Participant&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Each role has its own responsibilities and boundaries.&lt;/p&gt;

&lt;p&gt;We didn't want authorization to be something that existed only in the frontend.&lt;/p&gt;

&lt;p&gt;A judge shouldn't simply &lt;em&gt;not see&lt;/em&gt; something.&lt;/p&gt;

&lt;p&gt;The system should actually prevent them from accessing things they are not supposed to access.&lt;/p&gt;

&lt;p&gt;That distinction became increasingly important as judging logic became more complex.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhzawvcw0bspt852u96v9.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhzawvcw0bspt852u96v9.jpeg" alt="Role-aware platform administration and user management." width="800" height="374"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  3. We Treated Track and Non-Track Judging as Different Algorithmic Problems
&lt;/h2&gt;

&lt;p&gt;Instead of forcing every event through the same ranking and assignment flow, we explicitly considered:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Global submission pool&lt;/strong&gt; for non-track events.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Independent track pools&lt;/strong&gt; for multi-track events.&lt;/p&gt;

&lt;p&gt;That made judge eligibility and assignment much more predictable.&lt;/p&gt;


&lt;h2&gt;
  
  
  4. We Connected Assignment With Calibration
&lt;/h2&gt;

&lt;p&gt;This was probably our most important architectural insight.&lt;/p&gt;

&lt;p&gt;At first glance, these look like separate problems:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Assign judges.&lt;/p&gt;

&lt;p&gt;Normalize scores.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In reality, they are connected.&lt;/p&gt;

&lt;p&gt;The assignment strategy determines which judges overlap on which submissions.&lt;/p&gt;

&lt;p&gt;That overlap determines how much information exists for calibration.&lt;/p&gt;

&lt;p&gt;So the first algorithm influences the quality of the second.&lt;/p&gt;


&lt;h2&gt;
  
  
  Then Docker Happened.
&lt;/h2&gt;

&lt;p&gt;Not the interesting kind of happened.&lt;/p&gt;

&lt;p&gt;The painful kind.&lt;/p&gt;

&lt;p&gt;Our application had to run reliably through Docker, including under the constraints of the competition environment.&lt;/p&gt;

&lt;p&gt;The initial Docker setup itself wasn't particularly complicated.&lt;/p&gt;

&lt;p&gt;Then we discovered why our builds were taking absurdly long.&lt;/p&gt;
&lt;h2&gt;
  
  
  20–30 Minutes.
&lt;/h2&gt;

&lt;p&gt;For a build.&lt;/p&gt;

&lt;p&gt;We started looking at the application.&lt;/p&gt;

&lt;p&gt;Then the dependencies.&lt;/p&gt;

&lt;p&gt;Then Docker itself.&lt;/p&gt;

&lt;p&gt;Then everything else.&lt;/p&gt;

&lt;p&gt;The actual problem was embarrassingly simple.&lt;/p&gt;

&lt;p&gt;We had a &lt;strong&gt;large testing ZIP file sitting inside the project&lt;/strong&gt;, and we had forgotten to exclude it through &lt;code&gt;.dockerignore&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Docker was sending that file as part of the build context.&lt;/p&gt;

&lt;p&gt;So every build was dragging around a file that had absolutely no reason to be there.&lt;/p&gt;

&lt;p&gt;One missing &lt;code&gt;.dockerignore&lt;/code&gt; entry turned a normal development cycle into:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Make change
    ↓
docker compose up
    ↓
Wait
    ↓
Wait
    ↓
Wait
    ↓
Wonder what is happening
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once we found it, the fix was tiny.&lt;/p&gt;

&lt;p&gt;The debugging time wasn't.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Docker Lesson Wasn't "Use &lt;code&gt;.dockerignore&lt;/code&gt;"
&lt;/h2&gt;

&lt;p&gt;That would be too easy.&lt;/p&gt;

&lt;p&gt;The actual lesson was:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Build context is part of your build system.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When a Docker build is slow, it is tempting to immediately inspect dependencies, image layers, package installation, or application startup.&lt;/p&gt;

&lt;p&gt;Sometimes the problem happens before any of those things.&lt;/p&gt;

&lt;p&gt;The files you send to Docker matter.&lt;/p&gt;

&lt;p&gt;A massive build context can make the rest of your optimization efforts irrelevant.&lt;/p&gt;

&lt;p&gt;We learned that the hard way.&lt;/p&gt;




&lt;h2&gt;
  
  
  Another Weird Problem: Authentication Tests
&lt;/h2&gt;

&lt;p&gt;We also ran into inconsistent authentication test behaviour.&lt;/p&gt;

&lt;p&gt;The same &lt;code&gt;run.py&lt;/code&gt; tests could behave differently across systems.&lt;/p&gt;

&lt;p&gt;One system would pass authentication-related tests.&lt;/p&gt;

&lt;p&gt;Another would fail them.&lt;/p&gt;

&lt;p&gt;That created a particularly annoying debugging problem:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is the application broken?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is the environment different?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This was a useful reminder that reproducibility is not just about having the same source code.&lt;/p&gt;

&lt;p&gt;Environment state matters.&lt;/p&gt;

&lt;p&gt;Dependencies matter.&lt;/p&gt;

&lt;p&gt;Configuration matters.&lt;/p&gt;

&lt;p&gt;And when you're debugging during a hackathon, distinguishing between an application bug and an environment problem can consume an unreasonable amount of time.&lt;/p&gt;




&lt;h2&gt;
  
  
  What We Didn't Finish
&lt;/h2&gt;

&lt;p&gt;Not everything made it into the final implementation.&lt;/p&gt;

&lt;p&gt;Our proposed judging architecture included &lt;strong&gt;pairwise judging with the Bradley–Terry model&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We liked the idea because pairwise comparisons ask a different question.&lt;/p&gt;

&lt;p&gt;Instead of asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How many points does this project deserve?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;they ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Which of these two projects is better?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;From those comparisons, a ranking model can estimate relative strength.&lt;/p&gt;

&lt;p&gt;We had designed this as a potential judging stage.&lt;/p&gt;

&lt;p&gt;But we did not fully implement it as an organizer-selectable option within the available time.&lt;/p&gt;

&lt;p&gt;And we're actually glad we didn't rush it.&lt;/p&gt;

&lt;p&gt;A half-working statistical model buried inside the judging pipeline would have been much worse than explicitly leaving it unfinished.&lt;/p&gt;




&lt;h2&gt;
  
  
  The UI Lost the Race Against the Algorithm
&lt;/h2&gt;

&lt;p&gt;There was another tradeoff we had to make.&lt;/p&gt;

&lt;p&gt;We wanted a more polished UI.&lt;/p&gt;

&lt;p&gt;We didn't have unlimited time.&lt;/p&gt;

&lt;p&gt;Eventually we had to decide:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do we spend another few hours polishing visual details, or make sure the judging engine actually behaves correctly?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;We chose the latter.&lt;/p&gt;

&lt;p&gt;That meant sacrificing some visual polish in favour of getting the core judging workflow right.&lt;/p&gt;

&lt;p&gt;The interface is usable.&lt;/p&gt;

&lt;p&gt;It isn't the most visually sophisticated part of the project.&lt;/p&gt;

&lt;p&gt;But the system underneath it is much more interesting than the pixels suggest.&lt;/p&gt;

&lt;p&gt;And for this project, that felt like the right tradeoff.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Biggest Thing We Learned
&lt;/h2&gt;

&lt;p&gt;Before building this, we thought a hackathon platform was mainly a collection of features.&lt;/p&gt;

&lt;p&gt;Authentication.&lt;/p&gt;

&lt;p&gt;Submissions.&lt;/p&gt;

&lt;p&gt;Judges.&lt;/p&gt;

&lt;p&gt;Scores.&lt;/p&gt;

&lt;p&gt;Leaderboard.&lt;/p&gt;

&lt;p&gt;After building the judging system, we think about it differently.&lt;/p&gt;

&lt;p&gt;The difficult part isn't storing a score.&lt;/p&gt;

&lt;p&gt;It's deciding &lt;strong&gt;how that score should be produced, who should produce it, what constraints apply, how different evaluators affect it, and how the final result should be constructed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The leaderboard is just the final visible number.&lt;/p&gt;

&lt;p&gt;Underneath it is a chain of decisions.&lt;/p&gt;

&lt;p&gt;And every decision can introduce bias, inconsistency, or failure.&lt;/p&gt;

&lt;p&gt;That changed the way we approached the entire system.&lt;/p&gt;




&lt;h2&gt;
  
  
  What We Would Do Differently Next Time
&lt;/h2&gt;

&lt;p&gt;We would spend more time on the infrastructure around the algorithm before starting implementation.&lt;/p&gt;

&lt;p&gt;We would make the Docker build context explicit from day one.&lt;/p&gt;

&lt;p&gt;We would test authentication in more controlled environments earlier.&lt;/p&gt;

&lt;p&gt;We would build stronger test cases around judge assignment edge cases.&lt;/p&gt;

&lt;p&gt;And most importantly, we would formalize the judging model before writing too much application code.&lt;/p&gt;

&lt;p&gt;Because the hardest bugs weren't UI bugs.&lt;/p&gt;

&lt;p&gt;They weren't even ordinary backend bugs.&lt;/p&gt;

&lt;p&gt;They were &lt;strong&gt;modeling bugs&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We were trying to turn a real-world judging process—with constraints, people, preferences, fairness, and incomplete information—into deterministic software.&lt;/p&gt;

&lt;p&gt;That is a much harder problem than it first appears.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Architecture We Ended Up With
&lt;/h2&gt;

&lt;p&gt;At a high level, our judging pipeline became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Event
  ↓
Judging Configuration
  ↓
Track / Global Submission Pool
  ↓
Eligible Judges
  ↓
Feasibility Checks
  ↓
Deterministic Assignment
  ↓
Workload Balancing
  ↓
Overlap Optimization
  ↓
Judging
  ↓
WLS Calibration
  ↓
Ranking
  ↓
Validation
  ↓
Persisted Results
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part isn't that the pipeline looks neat.&lt;/p&gt;

&lt;p&gt;It's that each stage exists because we encountered a reason it needed to exist.&lt;/p&gt;




&lt;h2&gt;
  
  
  Final Takeaway
&lt;/h2&gt;

&lt;p&gt;We started with:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Let's build a hackathon platform."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;We ended up spending a surprising amount of time thinking about:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;allocation, fairness, statistical calibration, constraints, reproducibility, and failure modes.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And that's probably the biggest lesson we got from DOGFOOD 2026.&lt;/p&gt;

&lt;p&gt;Building the screens is the visible part of a hackathon platform.&lt;/p&gt;

&lt;p&gt;The difficult part is deciding what should happen &lt;strong&gt;after everyone clicks Submit&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Because once the submissions are in, the system has to make decisions.&lt;/p&gt;

&lt;p&gt;And those decisions need to be:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;deterministic enough to trust, flexible enough to configure, and explainable enough to defend.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That was the real system we were trying to build.&lt;/p&gt;

</description>
      <category>dogfood</category>
      <category>nextjs</category>
      <category>hackathonraptors</category>
    </item>
  </channel>
</rss>
