<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sukumar K</title>
    <description>The latest articles on DEV Community by Sukumar K (@sukumar_k_835e871c425c18d).</description>
    <link>https://dev.to/sukumar_k_835e871c425c18d</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4155591%2F9e96fe1b-a1c9-4e53-b8ea-a3bb1ade88a5.png</url>
      <title>DEV Community: Sukumar K</title>
      <link>https://dev.to/sukumar_k_835e871c425c18d</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sukumar_k_835e871c425c18d"/>
    <language>en</language>
    <item>
      <title>The Scoring Bug We Caught Before We Shipped It</title>
      <dc:creator>Sukumar K</dc:creator>
      <pubDate>Thu, 01 Oct 2026 18:27:49 +0000</pubDate>
      <link>https://dev.to/sukumar_k_835e871c425c18d/the-scoring-bug-we-caught-before-we-shipped-it-4enn</link>
      <guid>https://dev.to/sukumar_k_835e871c425c18d/the-scoring-bug-we-caught-before-we-shipped-it-4enn</guid>
      <description>&lt;p&gt;&lt;em&gt;Building ZenZone for DOGFOOD 2026 taught us that a fallback value can look reasonable and still be mathematically wrong.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Our judging system needed to account for judges who use the scoring scale differently. One judge might give nearly every project a 4; another might use the full range. We chose to normalize each judge’s scores onto a T-score scale:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;T = 50 + 10Z&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;That worked until we considered a judge who gave &lt;strong&gt;every project the same score&lt;/strong&gt;. Their scores have a standard deviation of zero, so the usual Z-score calculation would divide by zero. We needed a fallback.&lt;/p&gt;

&lt;p&gt;Our initial plan, recorded in &lt;code&gt;implementation_plan.md&lt;/code&gt;, said to substitute the event’s global mean score and log an audit record. It sounded sensible: if a judge provides no way to distinguish the projects they reviewed, use the average. But the plan missed one question: &lt;em&gt;the average on which scale?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Judges enter raw rubric scores. The value we store for the normalized result is a T-score centered on 50. The event’s global raw mean belongs to the first scale, not the second. Averaging it with other judges’ T-scores would mix unlike values.&lt;/p&gt;

&lt;p&gt;Consider a hypothetical project that receives a T-score of 60 from each of two judges. Suppose the event’s global &lt;strong&gt;raw&lt;/strong&gt; mean is 3.33. If we used that raw mean as the flat judge’s normalized score, the project’s combined result would be:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;(60 + 60 + 3.33) / 3 = 41.11&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Two judges placed the project above average, yet the fallback would pull its result below the T-score center of 50. That is not a judgment about the project; it is a scale mismatch.&lt;/p&gt;

&lt;p&gt;The correct neutral value in this calculation is &lt;strong&gt;50&lt;/strong&gt;. A judge who gives every project the same score provides no differential signal, which we represent as &lt;code&gt;Z = 0&lt;/code&gt;. Under our T-score formula, that becomes &lt;code&gt;T = 50&lt;/code&gt;. In the same hypothetical example:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;(60 + 60 + 50) / 3 = 56.67&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The flat judge no longer introduces an artificial penalty.&lt;/p&gt;

&lt;p&gt;What makes this a useful build story is where we caught it. We did &lt;strong&gt;not&lt;/strong&gt; ship a broken global-mean fallback and then discover bad production results. The written plan proposed one approach, but the committed implementation assigns &lt;code&gt;50.0&lt;/code&gt; when a judge’s score variance is effectively zero and writes a &lt;code&gt;ZERO_VARIANCE_FALLBACK&lt;/code&gt; audit entry.&lt;/p&gt;

&lt;p&gt;The code in &lt;code&gt;backend/src/main/java/com/dogfood/normalization/ZScoreNormalizationService.java&lt;/code&gt; still contains traces of the earlier plan: a comment describing “global mean substitution” and a calculation of &lt;code&gt;globalMean&lt;/code&gt; that the fallback no longer uses. Those leftovers are worth cleaning up. They also show why documentation and code need to be reviewed together: someone reading only that comment would come away with the wrong understanding of how ZenZone scores projects.&lt;/p&gt;

&lt;p&gt;The lesson I’m taking from this is broader than judging. &lt;strong&gt;A fallback has to be valid in the space where you use it.&lt;/strong&gt; Before substituting an average, default, or “neutral” value, check what that number represents—and whether every value in the final calculation is on the same scale.&lt;/p&gt;

</description>
      <category>bug</category>
      <category>debugging</category>
      <category>programming</category>
      <category>softwaredevelopment</category>
    </item>
  </channel>
</rss>
