You should freeze leaderboard tie-break rules as executable checks before any coding agent writes the ranking function. A polished draft can still reorder equal scores in ways your support replies never promised to players. This case study follows one small ranker from a written contract, through tests, to a review gate you can rerun. The conclusion you can keep is narrow: the tests define the order, and the agent does not.
Background you can recognize
You are shipping a public leaderboard for a small practice game that already stores one score row per player. Players compare screenshots, so the same two scores must produce the same order on every request. Your support notes will quote that order when someone asks why one name sits above another name. An agent can fill the function quickly, yet it still needs a rule you were willing to defend last week.
This week's public listings still spend visible attention on generated interfaces and on fast personal sites. That trend is useful surrounding context, but it does not decide how your leaderboard should break a tie. You can enjoy a fast generated draft and still owe players a stable order when two scores match exactly. The rest of this case study is the small project that makes that promised order executable and reviewable.
Goal of this small project
Your goal is a deterministic ranker that rejects incomplete rows instead of guessing a friendly order. You want the rule in a contract module, the examples in a test module, and the handler kept deliberately thin. You also want those checks to run in a clean environment, not only inside an editor session that hides local state. You are not measuring model speed, host capacity, or any contest outcome anywhere in this write-up.
Rules you freeze before prompting
Write the ranking rules in plain language before you paste a prompt into any coding assistant. A higher integer score always wins, and a missing score never outranks a row that stores a number. When two scores are equal, the earlier achieved_at value wins, and that value must include an explicit offset. When those timestamps also match, the smaller player_id string wins under ordinary Unicode code-point ordering.
What each case must do
- Complete rows that carry different scores sort by score descending, and the timestamp is not consulted.
- Complete rows that share one score sort by the earlier offset-aware timestamp, and then by player_id.
- A missing player_id, a missing score, or a naive timestamp raises ValueError before any list returns.
- The handler may shape JSON for clients, but it must not invent a second ordering rule of its own.
Decision table you keep beside the code
| Input situation | Score key | Time key | Id key | Expected result |
|---|---|---|---|---|
| Different scores | negative score | ignored | ignored | higher score first |
| Tied scores, different times | equal | achieved_at ascending | unused | earlier time first |
| Tied scores and times | equal | equal | player_id ascending | smaller id first |
| Missing field or naive time | not built | not built | not built | ValueError, no list |
You read this table before you accept a generated diff, because a green helper can still skip a required row. The negative score inside the key is only a sort trick so Python's ascending order matches your product order. The table remains the contract, and that trick is replaceable whenever the same tests still pass unchanged. If a later change needs another column, you update the table and the tests together in one commit.
Contract code you commit first
The listings below are an unexecuted example you can reproduce, not a log from a finished production deployment. You can copy the module into ranker.py and treat every raised error as part of the public behavior. You should not add hidden sort fields, such as signup date, unless the decision table already names them. If you prefer another language later, keep the same four outcomes and do not loosen those failures.
from dataclasses import dataclass
from datetime import datetime
from typing import Iterable
@dataclass(frozen=True)
class ScoreRow:
player_id: str
score: int
achieved_at: datetime
def sort_key(row: ScoreRow):
if not isinstance(row.player_id, str) or not row.player_id:
raise ValueError("player_id required")
if not isinstance(row.score, int) or isinstance(row.score, bool):
raise ValueError("score required")
if row.achieved_at.tzinfo is None or row.achieved_at.utcoffset() is None:
raise ValueError("achieved_at requires an offset")
return (-row.score, row.achieved_at, row.player_id)
def rank_rows(rows: Iterable[ScoreRow]) -> list[ScoreRow]:
return sorted(rows, key=sort_key)
Tests that reject a creative reorder
You write these tests before you ask for an implementation, and you commit them while the function is still unfinished. The fixture uses offset-aware timestamps so a naive datetime cannot sneak through as a quiet pass. You should keep the test names stable, because a review comment can cite the test instead of a chat log. After these tests exist, an implementation is acceptable only when this file remains untouched by the draft.
from datetime import datetime, timezone
import pytest
from ranker import ScoreRow, rank_rows
UTC = timezone.utc
EARLIER = datetime(2026, 10, 1, 12, 0, tzinfo=UTC)
LATER = datetime(2026, 10, 2, 12, 0, tzinfo=UTC)
def test_higher_score_wins_over_earlier_time():
rows = [
ScoreRow("ada", 10, EARLIER),
ScoreRow("bea", 12, LATER),
]
assert [row.player_id for row in rank_rows(rows)] == ["bea", "ada"]
def test_tie_prefers_earlier_timestamp_then_id():
rows = [
ScoreRow("cara", 10, LATER),
ScoreRow("bea", 10, EARLIER),
ScoreRow("ada", 10, EARLIER),
]
assert [row.player_id for row in rank_rows(rows)] == ["ada", "bea", "cara"]
def test_naive_timestamp_fails_closed():
naive = datetime(2026, 10, 1, 12, 0)
with pytest.raises(ValueError, match="offset"):
rank_rows([ScoreRow("ada", 10, naive)])
def test_bool_score_is_not_an_int_score():
with pytest.raises(ValueError, match="score"):
rank_rows([ScoreRow("ada", True, EARLIER)])
You run the module from a clean virtual environment so a global plugin cannot hide an import error. You expect four passed tests, and you treat a skipped test as a failed review rather than a partial success. If the third test fails, you inspect timezone handling before you change the sort key or the fixtures. If the fourth test fails, you remember that bool subclasses int, so a bare isinstance check is not enough.
python -m venv .venv
. .venv/bin/activate
python -m pip install pytest
python -m pytest tests/test_ranker.py -q
python - <<'PY'
from datetime import datetime, timezone
from ranker import ScoreRow, rank_rows
utc = timezone.utc
rows = [
ScoreRow("cara", 10, datetime(2026, 10, 2, 12, tzinfo=utc)),
ScoreRow("ada", 10, datetime(2026, 10, 1, 12, tzinfo=utc)),
ScoreRow("bea", 11, datetime(2026, 10, 3, 12, tzinfo=utc)),
]
print([row.player_id for row in rank_rows(rows)])
PY
You run that heredoc in bash, and on another shell you save the same lines as a file named check_order.py before you execute it. You should see bea, then ada, then cara, because the higher score outranks the tie and the earlier time selects ada. If you see cara before ada, the timestamp key is reversed and the second test should already be failing. You keep this driver out of production paths, because it is only a visibility check for the same frozen rule.
Implementation sequence
- You create the repository layout with ranker.py, a tests directory, and the decision table stored in the README.
- You commit the tests first, and you leave the ranker unfinished until those four outcomes are expressed clearly.
- You ask for a draft that edits ranker.py alone, and you rerun pytest before you open the generated diff.
- You merge only when the test file hash is unchanged and the handler does not sort the rows a second time.
Where a drafting tool fits this gate
A free model session is useful here only as a drafting surface for the unfinished ranker, after the contract is already in git. A free server option is useful when you want the same pytest command on a clean box you did not configure by hand. Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode is the open-source project this case study names when you want that optional drafting pass.
Free model access and a free server option are the only availability claims this article is willing to rely on. You should confirm the current terms yourself, and you should ignore any quota or hardware figure this page never states. The ranker, the table, and the tests remain useful if you delete every product name and draft the function yourself.
Thin handler you add only after green tests
import json
from ranker import rank_rows
def leaderboard_payload(rows):
ranked = rank_rows(rows)
body = {
"order": "score_desc_then_time_asc_then_id_asc",
"entries": [
{
"rank": index + 1,
"player_id": row.player_id,
"score": row.score,
"achieved_at": row.achieved_at.isoformat(),
}
for index, row in enumerate(ranked)
],
}
return json.dumps(body)
You add this wrapper only after the four tests pass, because a pretty payload can disguise a wrong order. The order string is a client label, not a second parser, so you keep it identical to the table heading. You do not sort again inside the comprehension, because a second sort is how silent disagreements return. If you expose this over HTTP later, you map ValueError to a client error and you log only the exception type.
Results you may report
You may report that the four tests encode the table, and that a passing run matches those fixtures exactly. You may not report a production error rate, a retention change, or a comparison against another assistant. This article includes no timed benchmark, server size, or token budget, so you should not invent those figures later. A passing local run is a review result for this fixture set, not proof about every future score row.
Lessons from the gate
The first lesson is that tie-breaks are product copy, so you freeze them before you enjoy a fast draft. The second lesson is that Python's bool subclass can smuggle a false score past a careless integer check. The third lesson is that timezone-naive timestamps feel convenient in tests and then disagree across machines. The fourth lesson is that an edited test is a contract change, even when the summary says the bug is fixed.
Limits you should say out loud
- This ranker does not handle pagination, cheating flags, or ties that appear only after display rounding.
- Equal timestamps across players are uncommon, but the id key keeps the API from depending on dict order.
- A free drafting session can still leak sample data if you paste real emails, tokens, or private user ids.
- Free model access and a free server can change, and this case study does not promise either one will last.
Who should not use this approach
You should skip this approach when real ordering depends on secret anti-cheat signals you must not place in a prompt. You should also skip it when a lawyer, not a test file, must define a prize ladder or a regional contest rule. You should skip it if nobody on your team will read the diff before merge, because tests do not replace that reading. You should skip any launch plan that assumes today's free access will still exist unchanged on your release day.
A closing check before you merge
Before you merge, you rerun the four tests on the edited commit and you confirm the test file hash stayed unchanged. You glance at the decision table, and you check that the handler only calls rank_rows before it builds JSON. You keep sample rows fictional, and you keep live credentials out of both the prompt and the server environment. If you want the same contract-first pass, read the current MonkeyCode docs before you try the free options.
Top comments (0)