<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: kadhiravan</title>
    <description>The latest articles on DEV Community by kadhiravan (@joker53).</description>
    <link>https://dev.to/joker53</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147379%2F512888d3-3d56-4252-a397-a14474bdbf5a.jpeg</url>
      <title>DEV Community: kadhiravan</title>
      <link>https://dev.to/joker53</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/joker53"/>
    <language>en</language>
    <item>
      <title>BeyondBug: The Score That Moved, the Boundary That Held</title>
      <dc:creator>kadhiravan</dc:creator>
      <pubDate>Tue, 29 Sep 2026 13:35:02 +0000</pubDate>
      <link>https://dev.to/joker53/beyondbug-the-score-that-moved-the-boundary-that-held-3kk</link>
      <guid>https://dev.to/joker53/beyondbug-the-score-that-moved-the-boundary-that-held-3kk</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;We built a self-hosted hackathon platform in 72 hours. The interface was the&lt;br&gt;
visible part. The real work was making deadlines, roles, judging, calibration,&lt;br&gt;
abuse controls, audit history, and offline operation agree with each other.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Three days is enough time to build a convincing interface. It is much less&lt;br&gt;
time than it sounds like when the interface must also enforce its own rules&lt;br&gt;
against direct HTTP requests, separate five roles, calibrate judges without&lt;br&gt;
hiding the original scores, and boot on an offline laptop from one command.&lt;/p&gt;

&lt;p&gt;BeyondBug is our answer to DOGFOOD 2026: an MIT-licensed submission and judging&lt;br&gt;
platform that takes an event from setup through registration, teams,&lt;br&gt;
submissions, judging, community voting, publication, feedback, awards, and&lt;br&gt;
certificates. This is a technical account of the decisions behind it, including&lt;br&gt;
the model we rejected, the model we eventually integrated, the authorization&lt;br&gt;
bug we found late, and the features we deliberately cut.&lt;/p&gt;

&lt;p&gt;Repository: &lt;a href="https://github.com/BeyondBug/DogFood" rel="noopener noreferrer"&gt;https://github.com/BeyondBug/DogFood&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  What shipped
&lt;/h2&gt;

&lt;p&gt;The product has five distinct actors:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Actor&lt;/th&gt;
&lt;th&gt;What the running backend permits&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Visitor&lt;/td&gt;
&lt;td&gt;Browse events, search and filter submitted projects, read published results and visible comments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Participant&lt;/td&gt;
&lt;td&gt;Register, form or join a team, save and preview a project, submit before the deadline, vote when eligible, and read their published feedback&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Judge&lt;/td&gt;
&lt;td&gt;Open only assigned projects, autosave a scorecard, submit a review, report a conflict, and read only their own scores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Organizer&lt;/td&gt;
&lt;td&gt;Configure an event, rubric, judge pool and assignments; inspect progress, audit history, calibration and voting; publish results and export data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Administrator&lt;/td&gt;
&lt;td&gt;Create events, provision organizers and judges, inspect system health and backups, and configure certificate designs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    V[Visitor] --&amp;gt; PUB[Public gallery and published results]
    P[Participant] --&amp;gt; SESSION[Session authentication]
    J[Judge] --&amp;gt; SESSION
    O[Organizer] --&amp;gt; SESSION
    A[Administrator] --&amp;gt; SESSION
    SESSION --&amp;gt; ROLE{Backend role and ownership checks}
    ROLE --&amp;gt;|participant and own team| PART[Team, submission, ballot, feedback]
    ROLE --&amp;gt;|assigned judge only| JUDGE[Own assignments and scorecards]
    ROLE --&amp;gt;|event organizer| ORG[Rubric, progress, audit, ranking, publication]
    ROLE --&amp;gt;|global administrator| ADMIN[Events, accounts, health, backups, designs]
    ROLE --&amp;gt;|scope mismatch| DENY[401 or 403]
    PUB --&amp;gt; API[FastAPI routes]
    PART --&amp;gt; API
    JUDGE --&amp;gt; API
    ORG --&amp;gt; API
    ADMIN --&amp;gt; API
    API --&amp;gt; TX[SQLite transaction]
    TX --&amp;gt; DATA[(Domain records)]
    TX --&amp;gt; AUDIT[(Audit entry)]

    classDef actor fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px;
    classDef public fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px;
    classDef gate fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:3px;
    classDef permitted fill:#ede9fe,stroke:#7c3aed,color:#2e1065,stroke-width:2px;
    classDef denied fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:3px;
    classDef system fill:#e2e8f0,stroke:#475569,color:#0f172a,stroke-width:2px;
    class V,P,J,O,A actor;
    class PUB public;
    class SESSION,ROLE gate;
    class PART,JUDGE,ORG,ADMIN permitted;
    class DENY denied;
    class API,TX,DATA,AUDIT system;&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The diagram's red path matters as much as its permitted paths: authorization&lt;br&gt;
is a backend decision made before protected records are read or changed.&lt;/p&gt;

&lt;p&gt;The browser experience includes role-specific dashboards, a four-step event&lt;br&gt;
setup flow, deadline and lifecycle indicators, submission readiness checks,&lt;br&gt;
judge autosave, next-project navigation, a readable audit log, light and dark&lt;br&gt;
modes, three visual themes, and responsive views. Those screens are backed by&lt;br&gt;
the same API used by the acceptance checker; none of the important boundaries&lt;br&gt;
depend on a hidden button.&lt;/p&gt;

&lt;p&gt;Projects support a repository, interactive demo, live URL, hosted video,&lt;br&gt;
thumbnail, gallery images and technology tags. Organizers can also define up&lt;br&gt;
to ten event-specific questions and require selected answers on final&lt;br&gt;
submission. Drafts may remain incomplete; the server validates required&lt;br&gt;
answers when a team submits. Assigned judges see those answers with the&lt;br&gt;
project, while unrelated accounts cannot read them.&lt;/p&gt;

&lt;p&gt;Organizers can export projects,&lt;br&gt;
participants, teams, judges, assignments, criterion-level scores, raw and&lt;br&gt;
adjusted rankings, audit history, votes, and certificates as CSV. Text fields&lt;br&gt;
are escaped against spreadsheet-formula injection.&lt;/p&gt;

&lt;p&gt;After publication, a participant sees anonymized criterion scores and written&lt;br&gt;
comments for their own team's project. Reviewer identities stay private. An&lt;br&gt;
organizer can assign configured prizes and issue separate participation and&lt;br&gt;
winner certificates. Each issued certificate stores a snapshot of its design,&lt;br&gt;
has a public database-backed verification page, and remains visually stable if&lt;br&gt;
an administrator later changes the template.&lt;/p&gt;

&lt;p&gt;The default local evaluation build also offers a demo-only role launcher. A&lt;br&gt;
reviewer can open the Administrator, Organizer, Judge A, Judge B, or&lt;br&gt;
Participant dashboard with one click. These are real expiring sessions using&lt;br&gt;
the same authorization checks, rather than frontend previews. Setting&lt;br&gt;
&lt;code&gt;DOGFOOD_DEMO_MODE=0&lt;/code&gt; removes the controls, makes the endpoint return 404, and&lt;br&gt;
revokes the fixed demo credentials on the existing volume.&lt;/p&gt;
&lt;h2&gt;
  
  
  We designed the denial paths first
&lt;/h2&gt;

&lt;p&gt;The most important request in BeyondBug returns no data:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Judge B → GET Judge A scores → 403 Forbidden
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The obvious implementation would filter scorecards in the page template. That&lt;br&gt;
would still let a judge change a query parameter or call the endpoint with&lt;br&gt;
&lt;code&gt;curl&lt;/code&gt;. We instead made the server establish identity from the session and&lt;br&gt;
check the requested judge ID against that identity. Participant requests to&lt;br&gt;
the same route receive 403. Rankings and exports use a separate organizer&lt;br&gt;
check.&lt;/p&gt;

&lt;p&gt;That choice shaped the schema. A user is global, while participant, judge and&lt;br&gt;
organizer roles belong to an event. A judge assignment links one accepted&lt;br&gt;
judge profile to one submitted project. A scorecard belongs to that assignment&lt;br&gt;
and preserves the rubric version. The write route can answer “does this&lt;br&gt;
session own this assignment?” without trusting a user ID supplied by the&lt;br&gt;
browser.&lt;/p&gt;

&lt;p&gt;Sessions use random opaque tokens; SQLite stores only SHA-256 digests.&lt;br&gt;
Passwords use salted PBKDF2-HMAC-SHA256 with 260,000 rounds. Cookies are&lt;br&gt;
HttpOnly and SameSite=Strict, can be marked Secure behind HTTPS, expire, and&lt;br&gt;
are revoked on logout or password change. A write carrying a foreign &lt;code&gt;Origin&lt;/code&gt;&lt;br&gt;
is rejected. Persistent login throttling blocks after five failed attempts for&lt;br&gt;
an account or twenty for a keyed client-IP digest in ten minutes. Unknown&lt;br&gt;
accounts still perform a dummy password check.&lt;/p&gt;

&lt;p&gt;We still found a role-model bug late in the build. A normal participant saw an&lt;br&gt;
event-creation form because the dashboard treated every authenticated account&lt;br&gt;
too similarly. Fixing the template alone would have repeated the original&lt;br&gt;
mistake. We made event creation administrator-only in both the UI and API, then&lt;br&gt;
added regression checks that participant, judge and organizer accounts see no&lt;br&gt;
creation form and receive HTTP 403 from the endpoint. Administrators can&lt;br&gt;
provision separate organizer accounts instead of sharing global authority.&lt;/p&gt;
&lt;h2&gt;
  
  
  Deadlines are database decisions
&lt;/h2&gt;

&lt;p&gt;The server uses UTC and checks a submission deadline inside the same&lt;br&gt;
&lt;code&gt;BEGIN IMMEDIATE&lt;/code&gt; transaction that changes the project or team. A stale page,&lt;br&gt;
modified browser clock or replayed request cannot reopen the event. The&lt;br&gt;
official fixture remains closed because its own &lt;code&gt;submissions_close&lt;/code&gt; timestamp&lt;br&gt;
is imported unchanged.&lt;/p&gt;

&lt;p&gt;Team creation, invitations and membership changes also stop at submission&lt;br&gt;
close. This matters to judging: a team roster used for conflict checks cannot&lt;br&gt;
change after assignments begin. Team invitation tokens are hashed, expire,&lt;br&gt;
work once, and preserve the four-person limit under a write lock.&lt;/p&gt;

&lt;p&gt;Publication is another state boundary. It requires closed submissions and a&lt;br&gt;
completed review for every ranked project. If community voting is configured,&lt;br&gt;
publication also waits for that window to close. Once results are public,&lt;br&gt;
projects, assignments and scorecards lock so the visible result cannot drift.&lt;/p&gt;
&lt;h2&gt;
  
  
  Assignment before arithmetic
&lt;/h2&gt;

&lt;p&gt;Normalization cannot repair a bad assignment graph. BeyondBug first assigns&lt;br&gt;
each submitted, nonduplicate project only to accepted judges who:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;selected the project's track;&lt;/li&gt;
&lt;li&gt;are not members of its team;&lt;/li&gt;
&lt;li&gt;have no declared conflict with the project; and&lt;/li&gt;
&lt;li&gt;are not already assigned to it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The allocator chooses the eligible judge with the lowest current load and&lt;br&gt;
uses stable judge ID as its tie-break. The target review count is configurable&lt;br&gt;
and defaults to three. If it cannot cover a project, it returns a visible&lt;br&gt;
shortage instead of silently assigning an ineligible judge. Each generated&lt;br&gt;
assignment and batch operation enters the audit trail.&lt;/p&gt;

&lt;p&gt;A judge can report a conflict only for their own unsubmitted assignment. The&lt;br&gt;
transaction records the conflict, removes the assignment and any draft&lt;br&gt;
scorecard, and prevents the same pair from being selected by a later batch.&lt;br&gt;
The organizer sees the new coverage gap. A conflict after final submission is&lt;br&gt;
not quietly rewritten; it returns 409 and requires human resolution.&lt;/p&gt;

&lt;p&gt;This is a deterministic load-balancing heuristic, not an optimal matching&lt;br&gt;
solver. It does not explicitly maximize overlap. The judging dashboard reports&lt;br&gt;
connected components in the judge/project graph because disconnected review&lt;br&gt;
pools cannot be calibrated against one another reliably.&lt;/p&gt;
&lt;h2&gt;
  
  
  Weighted scoring stays inspectable
&lt;/h2&gt;

&lt;p&gt;An organizer defines positive criterion weights before assignments exist.&lt;br&gt;
Judges score each criterion from 0 to 5. Draft scorecards may be partial;&lt;br&gt;
submission requires every criterion and rejects NaN, infinity and out-of-range&lt;br&gt;
values.&lt;/p&gt;

&lt;p&gt;For judge &lt;code&gt;j&lt;/code&gt;, project &lt;code&gt;p&lt;/code&gt; and criterion &lt;code&gt;c&lt;/code&gt;, the raw review score is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;raw(j,p) = Σ_c [weight(c) × 5 × score(j,p,c) / max_score(c)]
           --------------------------------------------------
                            Σ_c weight(c)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The shipped rubric uses &lt;code&gt;max_score(c)=5&lt;/code&gt;, so the result is a weighted average&lt;br&gt;
on the familiar five-point scale. Missing reviews remain missing rather than&lt;br&gt;
becoming zeroes. Review count is displayed beside every result, and projects&lt;br&gt;
without completed reviews remain unranked.&lt;/p&gt;
&lt;h2&gt;
  
  
  Normalization changed the winner
&lt;/h2&gt;

&lt;p&gt;An ordinary average assumes every judge uses the scale in the same way. Real&lt;br&gt;
panels contain strict and generous reviewers. BeyondBug fits a regularized&lt;br&gt;
two-way additive model over completed reviews of nonduplicate projects:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;raw(j,p) = quality(p) + severity(j) + error(j,p)

minimize:
    Σ_(j,p) [raw(j,p) - quality(p) - severity(j)]²
    + 3 × Σ_j severity(j)²
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The penalty of 3 shrinks judges with little evidence toward zero. Starting&lt;br&gt;
from each project's raw mean, the implementation alternates:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;severity(j) = Σ_p [raw(j,p) - quality(p)] / (review_count(j) + 3)
quality(p)  = mean_j [raw(j,p) - severity(j)]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Iteration stops when the maximum quality change is below &lt;code&gt;1e-10&lt;/code&gt;, or after&lt;br&gt;
500 iterations. Every review is adjusted with&lt;br&gt;
&lt;code&gt;clamp(raw - severity, 0, 5)&lt;/code&gt;, so calibration never leaves the rubric scale.&lt;br&gt;
The stored scorecard is never overwritten.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    RUBRIC[Weighted rubric] --&amp;gt; CARD[Submitted scorecards]
    CARD --&amp;gt; RAW[Raw 0-5 review scores]
    RAW --&amp;gt; GRAPH[Judge-project overlap graph]
    GRAPH --&amp;gt; CHECK{Connected review pool?}
    CHECK --&amp;gt;|No| WARN[Show comparison warning]
    CHECK --&amp;gt;|Yes| FIT[Fit regularized judge severity]
    FIT --&amp;gt; ADJUST[Clamp adjusted reviews to 0-5]
    ADJUST --&amp;gt; RANK[Adjusted project ranking]
    RAW --&amp;gt; PRESERVE[(Original scores preserved)]
    PRESERVE --&amp;gt; EXPLAIN[Raw vs adjusted explanation]
    RANK --&amp;gt; EXPLAIN
    WARN --&amp;gt; REVIEW[Organizer review]
    EXPLAIN --&amp;gt; REVIEW
    REVIEW --&amp;gt; PUBLISH{Publish results?}
    PUBLISH --&amp;gt;|Not ready| PRIVATE[Keep rankings private]
    PUBLISH --&amp;gt;|Approved after close| PUBLIC[Lock data and publish]

    classDef input fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px;
    classDef compute fill:#ede9fe,stroke:#7c3aed,color:#2e1065,stroke-width:2px;
    classDef evidence fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px;
    classDef decision fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:3px;
    classDef warning fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px;
    classDef outcome fill:#dcfce7,stroke:#16a34a,color:#14532d,stroke-width:3px;
    class RUBRIC,CARD,RAW input;
    class GRAPH,FIT,ADJUST,RANK compute;
    class PRESERVE,EXPLAIN,REVIEW evidence;
    class CHECK,PUBLISH decision;
    class WARN,PRIVATE warning;
    class PUBLIC outcome;&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Normalization is therefore an explanation pipeline, not a destructive rewrite:&lt;br&gt;
the organizer can always compare the original review with its adjustment.&lt;/p&gt;

&lt;p&gt;The official fixture gives this method a useful stress test:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;41 project records from 40 teams;&lt;/li&gt;
&lt;li&gt;one deliberate duplicate, &lt;code&gt;prj_41&lt;/code&gt;, with four historical reviews;&lt;/li&gt;
&lt;li&gt;126 historical scorecards in total;&lt;/li&gt;
&lt;li&gt;122 completed reviews over 40 ranked projects after excluding that duplicate;&lt;/li&gt;
&lt;li&gt;30 judges in one connected overlap component; and&lt;/li&gt;
&lt;li&gt;a constant-scoring judge whose score variance is zero.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because the estimator does not divide by a judge's standard deviation, the&lt;br&gt;
constant scorer stays finite. Missing batches also remain valid.&lt;/p&gt;

&lt;p&gt;The running calculation produces:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Project&lt;/th&gt;
&lt;th&gt;Raw rank&lt;/th&gt;
&lt;th&gt;Adjusted rank&lt;/th&gt;
&lt;th&gt;Adjusted score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Iron Switch&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;4.316&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Salt Ledger&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;4.295&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dry Relay&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;4.176&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Salt Loom&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;4.069&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Salt Kiln&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;4.043&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Thirty-three of the 40 ranked projects move. Open Beacon rises from 26 to 19;&lt;br&gt;
Paper Anchor falls from 21 to 28. The rank reversal is the reason this proof&lt;br&gt;
is interesting, but it is not proof that the adjusted order is objectively&lt;br&gt;
correct. It demonstrates that judge severity can affect an ordinary average&lt;br&gt;
and that our correction is reproducible.&lt;/p&gt;

&lt;p&gt;Regularization matters most when evidence is sparse. Judges &lt;code&gt;jdg_01&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;jdg_23&lt;/code&gt; have one review each. With the shipped penalty their offsets are&lt;br&gt;
about &lt;code&gt;-0.340&lt;/code&gt; and &lt;code&gt;-0.118&lt;/code&gt;; with a near-zero penalty of &lt;code&gt;0.01&lt;/code&gt;, they would&lt;br&gt;
be roughly &lt;code&gt;-1.804&lt;/code&gt; and &lt;code&gt;-0.992&lt;/code&gt;. The constant scorer &lt;code&gt;jdg_07&lt;/code&gt; receives a&lt;br&gt;
finite offset near &lt;code&gt;+0.209&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The organizer does not have to infer any of this from a CSV. The Judging&lt;br&gt;
Insight screen shows raw and adjusted score, rank movement, review coverage,&lt;br&gt;
overlap groups, judge offsets and the individual reviews behind a moved&lt;br&gt;
project. A positive displayed adjustment means the model identified a&lt;br&gt;
comparatively strict judge.&lt;/p&gt;
&lt;h2&gt;
  
  
  A deterministic attention signal comes first
&lt;/h2&gt;

&lt;p&gt;Before adding machine learning, we built a rule an organizer could audit. For&lt;br&gt;
each submitted review, BeyondBug compares its calibrated score with the median&lt;br&gt;
of the &lt;em&gt;other&lt;/em&gt; calibrated reviews on that project. It creates an attention&lt;br&gt;
item only when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;at least two peer reviews exist;&lt;/li&gt;
&lt;li&gt;the absolute gap is at least 1.5 points; and&lt;/li&gt;
&lt;li&gt;peer median absolute deviation is at most 0.5 points.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Projects with fewer than three reviews cannot trigger the signal. The detail&lt;br&gt;
page shows the original score, peer median, peer count, criterion values and&lt;br&gt;
written comment. Only organizers can open it. The rule never changes a score,&lt;br&gt;
assignment or award.&lt;/p&gt;

&lt;p&gt;That baseline gave us something crucial for the later ML integration: a clear&lt;br&gt;
fallback and a way to ask whether a model added useful prioritization rather&lt;br&gt;
than merely adding complexity.&lt;/p&gt;
&lt;h2&gt;
  
  
  We rejected the first ML model
&lt;/h2&gt;

&lt;p&gt;A teammate contributed an Isolation Forest for unusual reviews. Integrating a&lt;br&gt;
model quickly would have looked innovative, but its v1 contract did not match&lt;br&gt;
the product:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Problem&lt;/th&gt;
&lt;th&gt;v1 artifact&lt;/th&gt;
&lt;th&gt;Running portal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Score scale&lt;/td&gt;
&lt;td&gt;Synthetic 1–10 scores&lt;/td&gt;
&lt;td&gt;Weighted 0–5 rubric scores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inputs&lt;/td&gt;
&lt;td&gt;Included duration and edit count&lt;/td&gt;
&lt;td&gt;Neither field was recorded&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Peer statistics&lt;/td&gt;
&lt;td&gt;Included the review being evaluated&lt;/td&gt;
&lt;td&gt;Must exclude the candidate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Judge history&lt;/td&gt;
&lt;td&gt;Could include later reviews&lt;/td&gt;
&lt;td&gt;Live inference only knows earlier reviews&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Evaluation split&lt;/td&gt;
&lt;td&gt;Leaked reusable judge aggregates&lt;/td&gt;
&lt;td&gt;Needed complete event separation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime&lt;/td&gt;
&lt;td&gt;NumPy, joblib and scikit-learn pickle&lt;/td&gt;
&lt;td&gt;Offline image intentionally omitted them&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So v1 stayed in the repository as research history and did not enter the&lt;br&gt;
product. The integration review became a deployment gate: retrain on the right&lt;br&gt;
scale, remove leakage, report behavior by judge type, use only recorded fields,&lt;br&gt;
package inference offline, keep it organizer-only, and never let it make a&lt;br&gt;
scoring decision.&lt;/p&gt;
&lt;h2&gt;
  
  
  ML v2: a queue, never a verdict
&lt;/h2&gt;

&lt;p&gt;The retrained model uses a synthetic simulation designed around BeyondBug's&lt;br&gt;
actual contract:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Training fact&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Simulated events&lt;/td&gt;
&lt;td&gt;120&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Projects per event&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reviews per project&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total reviews&lt;/td&gt;
&lt;td&gt;14,400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Injected anomalies&lt;/td&gt;
&lt;td&gt;about 4.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Judge mix&lt;/td&gt;
&lt;td&gt;60% normal, 15% strict, 15% generous, 10% inconsistent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Isolation Forest&lt;/td&gt;
&lt;td&gt;300 trees, contamination 0.05&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Judge bias is stable within an event and judge IDs are unique across events.&lt;br&gt;
Peer features exclude the candidate review. Judge-history features use only&lt;br&gt;
earlier reviews. Train, validation and test data split by complete event; the&lt;br&gt;
reported test covers events 108–119. Duration and edit count were removed&lt;br&gt;
because the portal still does not measure them.&lt;/p&gt;

&lt;p&gt;The v2 model card reports these held-out synthetic results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Precision&lt;/td&gt;
&lt;td&gt;0.52&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recall&lt;/td&gt;
&lt;td&gt;0.56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;F1&lt;/td&gt;
&lt;td&gt;0.54&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Overall accuracy&lt;/td&gt;
&lt;td&gt;0.95&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decision-score gap&lt;/td&gt;
&lt;td&gt;0.137&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Accuracy is not the headline: anomalies are rare, so accuracy can look strong&lt;br&gt;
while the difficult class remains uncertain. Precision of 0.52 means many&lt;br&gt;
flags still need human judgment. The judge-type breakdown is more revealing:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Simulated judge type&lt;/th&gt;
&lt;th&gt;False-alarm rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Normal&lt;/td&gt;
&lt;td&gt;0.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inconsistent&lt;/td&gt;
&lt;td&gt;2.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Strict&lt;/td&gt;
&lt;td&gt;5.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generous&lt;/td&gt;
&lt;td&gt;7.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Consistently strict and generous judges are more likely to look unusual. We&lt;br&gt;
document that weakness instead of converting “unusual” into “dishonest.” The&lt;br&gt;
organizer sees High and Medium risk bands, the evidence count, and a link to&lt;br&gt;
the untouched scorecard. The official fixture yields 15 advisory signals; it&lt;br&gt;
has no anomaly labels, so that count is not reported as accuracy.&lt;/p&gt;

&lt;p&gt;The model is absent from every decision path. It cannot write a score, change&lt;br&gt;
normalization, assign a judge, disqualify anyone, select a winner, issue a&lt;br&gt;
certificate, or expose peer scores to judges. Judges and participants receive&lt;br&gt;
HTTP 403 from its endpoint.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    CONTRIB[Teammate model contribution] --&amp;gt; GATE{Integration gate}
    GATE --&amp;gt; SCALE[Matches 0-5 scale]
    GATE --&amp;gt; FEATURES[Uses recorded features]
    GATE --&amp;gt; LEAK[No target or future leakage]
    GATE --&amp;gt; SPLIT[Event-level evaluation split]
    GATE --&amp;gt; OFFLINE[Offline reviewed runtime]
    GATE --&amp;gt; AUTH[Organizer-only access]
    SCALE --&amp;gt; PASS{All checks pass?}
    FEATURES --&amp;gt; PASS
    LEAK --&amp;gt; PASS
    SPLIT --&amp;gt; PASS
    OFFLINE --&amp;gt; PASS
    AUTH --&amp;gt; PASS
    PASS --&amp;gt;|v1: no| RESEARCH[Keep as research history]
    PASS --&amp;gt;|v2: yes| JSON[Export 300 trees to compressed JSON]
    JSON --&amp;gt; SIGNAL[Advisory review signals]
    SIGNAL --&amp;gt; HUMAN[Organizer inspects evidence and scorecard]
    HUMAN --&amp;gt; DECISION[Human decision outside the model]
    SIGNAL -. never writes .-&amp;gt; PROTECTED[Scores, rankings, assignments, awards]

    classDef source fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px;
    classDef gate fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:3px;
    classDef check fill:#ede9fe,stroke:#7c3aed,color:#2e1065,stroke-width:2px;
    classDef rejected fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px;
    classDef accepted fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px;
    classDef human fill:#dcfce7,stroke:#16a34a,color:#14532d,stroke-width:3px;
    classDef protected fill:#f1f5f9,stroke:#475569,color:#0f172a,stroke-width:2px,stroke-dasharray:5 5;
    class CONTRIB source;
    class GATE,PASS gate;
    class SCALE,FEATURES,LEAK,SPLIT,OFFLINE,AUTH check;
    class RESEARCH rejected;
    class JSON,SIGNAL accepted;
    class HUMAN,DECISION human;
    class PROTECTED protected;&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The dashed edge is intentionally a non-effect: the signal can point an&lt;br&gt;
organizer toward evidence, but it has no write path into judging outcomes.&lt;/p&gt;
&lt;h2&gt;
  
  
  Pairwise mode stayed separate from the official ranking
&lt;/h2&gt;

&lt;p&gt;The primary result remains the weighted rubric with transparent judge-severity&lt;br&gt;
calibration. Pairwise Mode is a separate experiment: the organizer generates&lt;br&gt;
project pairs, assigned judges choose one project from each pair, and a&lt;br&gt;
regularized Bradley–Terry estimator recovers latent strengths. Stable ordering&lt;br&gt;
and assignment ownership keep the flow reproducible and isolated.&lt;/p&gt;

&lt;p&gt;A pairwise vote never edits a criterion score, normalized ranking, prize, or&lt;br&gt;
certificate. Organizers can compare both views and explain disagreement&lt;br&gt;
without silently replacing the published method. Sparse comparisons remain&lt;br&gt;
visible as limited evidence rather than being presented as certainty.&lt;/p&gt;

&lt;p&gt;The same isolated module provides CSV account import, a public gallery embed,&lt;br&gt;
a sanitized event archive, Ed25519-signed judge participation records, and&lt;br&gt;
HMAC-signed webhook deliveries with visible status. A signed judge record&lt;br&gt;
proves its payload matches the issuer's signature; an external verifier still&lt;br&gt;
needs to pin or otherwise trust that issuer key.&lt;/p&gt;
&lt;h2&gt;
  
  
  Shipping ML without shipping an ML runtime
&lt;/h2&gt;

&lt;p&gt;The one-command offline rule made model packaging as important as training.&lt;br&gt;
Loading the joblib file would require a compatible pickle environment plus&lt;br&gt;
NumPy and scikit-learn. Instead, we exported all 300 trees to compressed JSON&lt;br&gt;
and implemented the Isolation Forest decision function with the Python&lt;br&gt;
standard library.&lt;/p&gt;

&lt;p&gt;The runtime reads the ordered feature list exactly, including intentionally&lt;br&gt;
repeated names used as feature weighting. A regression test compares the&lt;br&gt;
portable decision score with the original scikit-learn artifact on a reference&lt;br&gt;
vector. The Docker image loads no pickle and installs no ML library. On the&lt;br&gt;
fixture, the organizer-only model panel rendered in about 0.13 seconds during&lt;br&gt;
release verification.&lt;/p&gt;

&lt;p&gt;This was the ML lesson of the build: integration quality includes the feature&lt;br&gt;
contract, leakage boundary, evaluation split, runtime format, authorization,&lt;br&gt;
UI wording and effect on downstream decisions. A model file alone is not a&lt;br&gt;
feature.&lt;/p&gt;
&lt;h2&gt;
  
  
  Community voting assumes attackers exist
&lt;/h2&gt;

&lt;p&gt;BeyondBug supports disabled voting, curated email-bound invitations, or&lt;br&gt;
authenticated event participants. Ballot projects are ordered by an HMAC of a&lt;br&gt;
secret event seed, voter ID and project ID. The order is stable when one voter&lt;br&gt;
refreshes but differs between voters, avoiding both insertion-order bias and a&lt;br&gt;
frustrating reshuffle on every page load.&lt;/p&gt;

&lt;p&gt;The database permits one final ballot per account and event. The API rejects&lt;br&gt;
self-votes, duplicate projects and flagged duplicate submissions. Voting&lt;br&gt;
eligibility and time windows are checked again inside the transaction. Vote&lt;br&gt;
attempts are limited by account and keyed IP digest; comments require login,&lt;br&gt;
reject exact repeats and are limited to five per account per hour. Organizers&lt;br&gt;
can hide comments without erasing the moderation record.&lt;/p&gt;

&lt;p&gt;Totals remain private through the voting window. Public result routes return&lt;br&gt;
404 until publication; the organizer can see interim signals without leaking&lt;br&gt;
them to voters. Voting configuration locks after the first ballot so an&lt;br&gt;
organizer cannot silently change eligibility midstream.&lt;/p&gt;

&lt;p&gt;This does not solve Sybil identity. Email matching does not prove inbox&lt;br&gt;
ownership, an account does not prove one human, and shared networks complicate&lt;br&gt;
IP limits. We recommend curated invitations for high-stakes community prizes&lt;br&gt;
and say so in the threat model.&lt;/p&gt;
&lt;h2&gt;
  
  
  Audit history has to be readable
&lt;/h2&gt;

&lt;p&gt;Writing rows called &lt;code&gt;project.updated&lt;/code&gt; and &lt;code&gt;scorecard.submitted&lt;/code&gt; satisfies a&lt;br&gt;
database requirement but does not help an operator during an event. The&lt;br&gt;
organizer view resolves actors and targets, presents plain-language actions,&lt;br&gt;
groups activity into event, team, submission, judging, voting, security and&lt;br&gt;
publication categories, and supports filtering.&lt;/p&gt;

&lt;p&gt;Consequential writes include the actor, entity, action, UTC timestamp and JSON&lt;br&gt;
details. The log covers event changes, invitations, team membership, project&lt;br&gt;
submission, assignment batches, conflicts, scorecards, voting configuration,&lt;br&gt;
ballots, comments, moderation, publication, prizes and certificate issuance.&lt;br&gt;
It is useful for diagnosis, but we do not call it tamper-proof: a malicious&lt;br&gt;
host operator who controls SQLite can change the log. Signed external&lt;br&gt;
transparency records remain future work.&lt;/p&gt;
&lt;h2&gt;
  
  
  Offline operation changed normal product decisions
&lt;/h2&gt;

&lt;p&gt;There is no hosted database, authentication service, email provider, CDN,&lt;br&gt;
analytics endpoint or external API. FastAPI, SQLite, fonts, templates, scripts,&lt;br&gt;
the model export, fixture data and pinned Python wheels live in the repository.&lt;br&gt;
The image supports x86-64 and ARM64 wheels and installs them with &lt;code&gt;--no-index&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker compose up
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On first boot, migrations run through schema version 10, the official fixture&lt;br&gt;
is imported once, and demo authorization headers are printed. The sign-in page&lt;br&gt;
also exposes five demo role shortcuts. Restarts preserve changes. Demo mode can&lt;br&gt;
be disabled for a real deployment, with the global administrator bootstrapped&lt;br&gt;
from local environment variables.&lt;/p&gt;

&lt;p&gt;SQLite runs with WAL, foreign keys and a ten-second busy timeout. Immediate&lt;br&gt;
write transactions protect team limits, deadlines, assignments and ballots.&lt;br&gt;
The online backup command creates a consistent snapshot while the portal is&lt;br&gt;
running. An administrator can also run an integrity check and download local&lt;br&gt;
snapshots from the system page. A stopped installation can validate and&lt;br&gt;
atomically restore a full SQLite snapshot while preserving the previous&lt;br&gt;
database as a safety copy.&lt;/p&gt;

&lt;p&gt;For migration between installations, an organizer can export a pre-judging&lt;br&gt;
portable JSON bundle containing tracks, prizes, custom questions, participants,&lt;br&gt;
teams, projects, answers, and judge profiles. Import first performs a dry run,&lt;br&gt;
requires an empty open target event, validates every reference, and commits in&lt;br&gt;
one transaction. Newly created local passwords are shown once. Historical&lt;br&gt;
reviews, ballots, certificates, and audit rows stay outside this portable&lt;br&gt;
format; a complete move uses the SQLite snapshot.&lt;/p&gt;

&lt;p&gt;We did not add a decorative load balancer. One Uvicorn worker and one SQLite&lt;br&gt;
database form the supported deployment. On a development laptop, a warm local&lt;br&gt;
read probe against the 41-project gallery measured:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requests&lt;/th&gt;
&lt;th&gt;Workers&lt;/th&gt;
&lt;th&gt;Success&lt;/th&gt;
&lt;th&gt;Throughput&lt;/th&gt;
&lt;th&gt;p50&lt;/th&gt;
&lt;th&gt;p95&lt;/th&gt;
&lt;th&gt;Maximum&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;500/500&lt;/td&gt;
&lt;td&gt;359.5 req/s&lt;/td&gt;
&lt;td&gt;55.1 ms&lt;/td&gt;
&lt;td&gt;63.4 ms&lt;/td&gt;
&lt;td&gt;70.5 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1,000&lt;/td&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;1,000/1,000&lt;/td&gt;
&lt;td&gt;336.6 req/s&lt;/td&gt;
&lt;td&gt;145.5 ms&lt;/td&gt;
&lt;td&gt;176.2 ms&lt;/td&gt;
&lt;td&gt;221.9 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are short read tests, not a production service-level objective and not&lt;br&gt;
a simultaneous-user rating. They do not measure write contention. If measured&lt;br&gt;
event traffic exceeds this design, the next architecture needs a shared&lt;br&gt;
database, shared rate limiting and background jobs before multiple application&lt;br&gt;
instances and a reverse proxy become meaningful.&lt;/p&gt;
&lt;h2&gt;
  
  
  Evidence over claims
&lt;/h2&gt;

&lt;p&gt;The required checker makes seven HTTP requests. The committed, unedited report&lt;br&gt;
passes all seven and verifies T1 and T2:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T1  gallery is public ................. PASS
T1  project from fixtures shown ....... PASS
T1  closed event refuses submissions .. PASS
T2  judge sees own scores ............. PASS
T2  judge cannot see peer scores ...... PASS
T2  participant blocked ............... PASS
T2  csv export works .................. PASS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Our separate test runner starts a disposable Compose project on a random port,&lt;br&gt;
creates a fresh named volume, runs 39 unit and HTTP integration tests, and&lt;br&gt;
removes only that test environment. It covers deadlines, role denials, team&lt;br&gt;
limits, conflicts, publication locks, normalization edge cases, stable ballot&lt;br&gt;
randomization, rate limits, duplicate handling, certificate behavior, backups,&lt;br&gt;
OpenAPI synchronization, portable ML inference, Pairwise Mode, webhooks,&lt;br&gt;
custom questions, portable bundle round trips, safe restore, bulk import,&lt;br&gt;
embeds, Ed25519 judge records, and all five demo role tours. The combined&lt;br&gt;
release result is 39/39.&lt;/p&gt;

&lt;p&gt;We also tested the current image with its Docker network disconnected;&lt;br&gt;
its local health endpoint returned HTTP 200. The five-minute lifecycle video&lt;br&gt;
shows real browser actions from event creation to publication, including the&lt;br&gt;
direct peer-score 403 and a CSV export.&lt;/p&gt;

&lt;p&gt;Our &lt;code&gt;.dogfood.toml&lt;/code&gt; still claims only T1 and T2 because those are the tiers the&lt;br&gt;
provided checker can verify. T3 voting and comments have their own requirement-&lt;br&gt;
to-test evidence document, but we do not label them checker-certified. Honest&lt;br&gt;
scope is more valuable than a larger label.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we cut
&lt;/h2&gt;

&lt;p&gt;We isolated Bradley–Terry pairwise judging from the primary rubric ranking so&lt;br&gt;
an experimental comparison mode cannot silently change official results. We&lt;br&gt;
also added signed webhooks, an embeddable gallery, Ed25519 judge participation&lt;br&gt;
records, bulk CSV account import, portable pre-judging exchange, and validated&lt;br&gt;
command-line restore. The remaining cuts are account recovery, email delivery,&lt;br&gt;
anonymous open-link voting, exhaustive webhook coverage, and browser-based&lt;br&gt;
restore.&lt;/p&gt;

&lt;p&gt;Certificates are publicly verifiable against the local database, but they are&lt;br&gt;
not cryptographically signed. Backups are local snapshots, not scheduled&lt;br&gt;
off-host disaster recovery. Duplicate submissions are detected by identical&lt;br&gt;
nonempty repository URLs; legitimate forks can be flagged and copied work at a&lt;br&gt;
different URL can be missed. Published score correction needs a future&lt;br&gt;
versioned republication workflow.&lt;/p&gt;

&lt;p&gt;Those limits appear in the README, architecture, data model and threat model.&lt;br&gt;
The objective was software another organizer could evaluate, operate and&lt;br&gt;
extend, not a checklist with hidden gaps.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we would redo
&lt;/h2&gt;

&lt;p&gt;We would model immutable result snapshots from the start. The current lock&lt;br&gt;
prevents rankings from drifting, but production events eventually need a&lt;br&gt;
correction, explanation and republication history. We would also include rich&lt;br&gt;
submission media and organizer-defined questions in the first schema instead&lt;br&gt;
of adding project media during migration 7.&lt;/p&gt;

&lt;p&gt;For judging, we would optimize assignment overlap explicitly rather than only&lt;br&gt;
balancing count among eligible judges. The current component warning makes&lt;br&gt;
disconnected pools visible, but prevention is stronger than diagnosis.&lt;/p&gt;

&lt;p&gt;For the ML work, the next useful step is evaluation on carefully adjudicated&lt;br&gt;
real-event data, with calibration by review count and judge type. Until such&lt;br&gt;
data exists, the model should remain an inspection queue with prominent limits.&lt;br&gt;
Adding duration or edit behavior would require privacy review and reliable&lt;br&gt;
instrumentation before retraining.&lt;/p&gt;

&lt;p&gt;The largest lesson was that fairness features need explanation surfaces.&lt;br&gt;
Normalization hidden in a backend function is hard to defend. A model badge&lt;br&gt;
without its evidence count looks accusatory. An audit table without human names&lt;br&gt;
and actions is operationally useless. BeyondBug became stronger whenever the&lt;br&gt;
system showed its reasoning, its raw inputs and its limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/BeyondBug/DogFood.git
&lt;span class="nb"&gt;cd &lt;/span&gt;DogFood
docker compose up
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then open &lt;code&gt;http://localhost:8080&lt;/code&gt;. The repository includes the fixture,&lt;br&gt;
acceptance checker, unedited acceptance report, OpenAPI document, architecture,&lt;br&gt;
data model, judging proof, threat model, ML model card, integration review,&lt;br&gt;
capacity probe, test runner and five-minute demo. In the default demo build,&lt;br&gt;
choose a role directly on the sign-in page. Production operators disable that&lt;br&gt;
launcher with &lt;code&gt;DOGFOOD_DEMO_MODE=0&lt;/code&gt; and provide bootstrap administrator&lt;br&gt;
credentials through local environment variables.&lt;/p&gt;

&lt;p&gt;BeyondBug is available under the MIT license. Built by team BeyondBug for&lt;/p&gt;

&lt;h1&gt;
  
  
  DogfoodHackathon.
&lt;/h1&gt;

</description>
      <category>dogfoodhackathon</category>
      <category>hackathon</category>
      <category>opensource</category>
      <category>security</category>
    </item>
  </channel>
</rss>
