DEV Community

Sukumar K
Sukumar K

Posted on

The Scoring Bug We Caught Before We Shipped It

Building ZenZone for DOGFOOD 2026 taught us that a fallback value can look reasonable and still be mathematically wrong.

Our judging system needed to account for judges who use the scoring scale differently. One judge might give nearly every project a 4; another might use the full range. We chose to normalize each judge’s scores onto a T-score scale:

T = 50 + 10Z

That worked until we considered a judge who gave every project the same score. Their scores have a standard deviation of zero, so the usual Z-score calculation would divide by zero. We needed a fallback.

Our initial plan, recorded in implementation_plan.md, said to substitute the event’s global mean score and log an audit record. It sounded sensible: if a judge provides no way to distinguish the projects they reviewed, use the average. But the plan missed one question: the average on which scale?

Judges enter raw rubric scores. The value we store for the normalized result is a T-score centered on 50. The event’s global raw mean belongs to the first scale, not the second. Averaging it with other judges’ T-scores would mix unlike values.

Consider a hypothetical project that receives a T-score of 60 from each of two judges. Suppose the event’s global raw mean is 3.33. If we used that raw mean as the flat judge’s normalized score, the project’s combined result would be:

(60 + 60 + 3.33) / 3 = 41.11

Two judges placed the project above average, yet the fallback would pull its result below the T-score center of 50. That is not a judgment about the project; it is a scale mismatch.

The correct neutral value in this calculation is 50. A judge who gives every project the same score provides no differential signal, which we represent as Z = 0. Under our T-score formula, that becomes T = 50. In the same hypothetical example:

(60 + 60 + 50) / 3 = 56.67

The flat judge no longer introduces an artificial penalty.

What makes this a useful build story is where we caught it. We did not ship a broken global-mean fallback and then discover bad production results. The written plan proposed one approach, but the committed implementation assigns 50.0 when a judge’s score variance is effectively zero and writes a ZERO_VARIANCE_FALLBACK audit entry.

The code in backend/src/main/java/com/dogfood/normalization/ZScoreNormalizationService.java still contains traces of the earlier plan: a comment describing “global mean substitution” and a calculation of globalMean that the fallback no longer uses. Those leftovers are worth cleaning up. They also show why documentation and code need to be reviewed together: someone reading only that comment would come away with the wrong understanding of how ZenZone scores projects.

The lesson I’m taking from this is broader than judging. A fallback has to be valid in the space where you use it. Before substituting an average, default, or “neutral” value, check what that number represents—and whether every value in the final calculation is on the same scale.

Top comments (0)