DEV Community

Alex Chen
Alex Chen

Posted on

Learn Holdout Leakage by Building a Tiny Split Receipt

Last Tuesday the rain was doing that Halifax thing where it looks polite and still soaks your notes. I was in the Killam stacks with a six-line CSV and a README sentence a draft had written for me: the holdout looked stable. Stable how?

The file had six comments and five student ids. One person had spoken twice. If I cut the file by row, that person could sit in both rooms and smile at my accuracy like a classmate who took the quiz twice.

That is the whole case. I was not training a real model yet. I was about to pin a tiny portfolio lab, and I wanted one proof that the holdout was actually a holdout.

Here is the fixture I should have stared at before I trusted the sentence. Read the first data line and the last one before you read any code.

student_id,text,label
s1,the lab wifi dropped again,neg
s2,pytorch install finally worked,pos
s3,i do not understand the loss,neg
s4,office hours saved me,pos
s5,the holdout looks too clean,neg
s1,same lab second comment,neg
Enter fullscreen mode Exit fullscreen mode

Same student_id. Different sentence. A row split does not care, and a person split does. Which one did the draft assume?

It never said. That silence is how a student project starts lying without a single syntax error.

The learning question is small enough to finish between classes. Can I force every id onto one side of the split, write a receipt I can re-run, and reject a claimed split that sneaks an id into both sides? If I cannot, I have no business putting a metric in a README.

I needed three things and nothing else. Python 3.11 or newer, the standard library, and a fixture I could read with my eyes. No pandas. No notebook, and no hidden shuffle.

If a step is not in the file, I cannot defend it in a lab report. That rule has saved me more often than any model has. You can call that stubborn. I call it the only way I remember what I did.

The goal was not a better classifier. The goal was a receipt: train ids, test ids, the overlap, and a short digest of that pair. If the overlap is not empty, the process exits non-zero.

Pretty output is not a pass. I would rather have an ugly exit code than a smooth paragraph.

I wrote one script and named it split_receipt.py. The grouped path sorts unique ids and parks a tail slice in test. Holdout size is max(1, (n * 2) // 5), an integer stand-in for about 40 percent, so I do not have to argue with binary floats. With five ids that is 2, so test gets s4 and s5, and train keeps s1, s2, and s3.

Both of s1's comments stay with s1. The naive path ignores identity and takes the last row as test. On this fixture the last row is s1 again, so the overlap is s1 and the script exits 1.

Same file. Two stories. Only one of them is a holdout.

I am not going to paste a fake accuracy next to either story. I never trained anything, so the result of this case is the exit code.

#!/usr/bin/env python3
"""Tiny holdout receipt. Python 3.11+. Standard library only."""

import csv
import hashlib
import json
import sys
from pathlib import Path


def load_rows(path: Path) -> list[dict[str, str]]:
    with path.open(newline="", encoding="utf-8") as handle:
        rows = list(csv.DictReader(handle))
    if not rows:
        raise SystemExit(f"empty fixture: {path}")
    return rows


def unique_ids(rows: list[dict[str, str]], id_col: str) -> list[str]:
    missing = [i for i, row in enumerate(rows) if not row.get(id_col, "").strip()]
    if missing:
        raise SystemExit(f"missing {id_col} on row indexes {missing}")
    return sorted({row[id_col].strip() for row in rows})


def assign_by_id(ids: list[str]) -> tuple[list[str], list[str]]:
    holdout_count = max(1, (len(ids) * 2) // 5)
    test_ids = ids[-holdout_count:]
    train_ids = ids[:-holdout_count]
    if not train_ids:
        raise SystemExit("holdout ate the training ids; use more ids")
    return train_ids, test_ids


def naive_by_row(
    rows: list[dict[str, str]], id_col: str, test_count: int = 1
) -> tuple[list[str], list[str]]:
    missing = [i for i, row in enumerate(rows) if not row.get(id_col, "").strip()]
    if missing:
        raise SystemExit(f"missing {id_col} on row indexes {missing}")
    if test_count >= len(rows):
        raise SystemExit("naive split left no training rows")
    train_ids = [row[id_col].strip() for row in rows[:-test_count]]
    test_ids = [row[id_col].strip() for row in rows[-test_count:]]
    return train_ids, test_ids


def build_receipt(train_ids: list[str], test_ids: list[str], id_col: str) -> dict:
    overlap = sorted(set(train_ids) & set(test_ids))
    payload = json.dumps(
        {"train": train_ids, "test": test_ids}, separators=(",", ":")
    )
    digest = hashlib.sha256(payload.encode("utf-8")).hexdigest()[:12]
    return {
        "id_col": id_col,
        "train_ids": train_ids,
        "test_ids": test_ids,
        "overlap": overlap,
        "ok": (not overlap) and bool(train_ids) and bool(test_ids),
        "digest": digest,
    }


def clean_row(row: dict[str, str]) -> dict[str, str]:
    needed = ("student_id", "text", "label")
    if any(key not in row for key in needed):
        raise ValueError("row missing student_id, text, or label")
    text = " ".join(row["text"].split()).lower()
    if not text:
        raise ValueError("empty text")
    label = row["label"].strip().lower()
    if label not in {"pos", "neg"}:
        raise ValueError(f"bad label: {label}")
    return {
        "student_id": row["student_id"].strip(),
        "text": text,
        "label": label,
    }


def check_claimed(path: Path, known_ids: list[str]) -> dict:
    data = json.loads(path.read_text(encoding="utf-8"))
    train = [str(item) for item in data.get("train_ids", [])]
    test = [str(item) for item in data.get("test_ids", [])]
    unknown = sorted((set(train) | set(test)) - set(known_ids))
    overlap = sorted(set(train) & set(test))
    ok = bool(train) and bool(test) and not unknown and not overlap
    return {"unknown": unknown, "overlap": overlap, "ok": ok}


def main() -> None:
    if len(sys.argv) < 2:
        raise SystemExit(
            "usage: python3 split_receipt.py comments.csv [--naive] [claimed.json]"
        )
    fixture = Path(sys.argv[1])
    flags = sys.argv[2:]
    naive = "--naive" in flags
    claimed_args = [item for item in flags if not item.startswith("--")]
    rows = load_rows(fixture)
    ids = unique_ids(rows, "student_id")
    if naive:
        train_ids, test_ids = naive_by_row(rows, "student_id", 1)
    else:
        train_ids, test_ids = assign_by_id(ids)
    receipt = build_receipt(train_ids, test_ids, "student_id")
    print(json.dumps(receipt, indent=2))
    if claimed_args:
        verdict = check_claimed(Path(claimed_args[0]), ids)
        print("claimed " + json.dumps(verdict))
        if not verdict["ok"]:
            raise SystemExit(2)
    if not receipt["ok"]:
        raise SystemExit(1)


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Save the CSV as comments.csv in the same directory. Then run the two commands below. The first is the split I am willing to show a classmate. The second is the cut I almost shipped.

python3 split_receipt.py comments.csv
python3 split_receipt.py comments.csv --naive
Enter fullscreen mode Exit fullscreen mode

The first command should print "ok": true, train ids s1 s2 s3, test ids s4 s5, and an empty overlap. The digest is the first 12 hex characters of SHA-256 over the canonical JSON of those two lists, with no spaces. I am not going to type that digest from memory. Run it twice, and if the two digests differ, you changed the split or the serializer.

The second command is the error input I actually care about. Overlap should be a list containing only s1, ok should be false, and the process should exit 1. If your naive run exits 0, you do not have my fixture.

Do not fix the checker to make it pass. Fix the file, or admit you wanted a row cut.

There is a second bad artifact, for the morning a draft hands you a split you did not compute. Save the JSON below as claimed_overlap.json. Before you run it, which id do you think is sitting in both lists?

{"train_ids": ["s1", "s2"], "test_ids": ["s2", "s5"]}
Enter fullscreen mode Exit fullscreen mode

I hope you said s2. Then run the checker with that file and no naive flag. You should see the good receipt first, then a claimed object whose overlap is s2 and whose ok is false, then exit code 2.

python3 split_receipt.py comments.csv claimed_overlap.json
Enter fullscreen mode Exit fullscreen mode

Unknown ids fail the same way. A draft that invents s9 is not being creative. It is describing a student who is not in the room.

Cleaning is the other place a draft likes to look finished. I kept a tiny function in the same file because empty text is a different lie from a leaked id. Whitespace is not a comment.

python3 -c 'from split_receipt import clean_row; clean_row({"student_id":"s9","text":"   ","label":"pos"})'
Enter fullscreen mode Exit fullscreen mode

That should raise ValueError: empty text. A label like maybe should raise bad label: maybe. I want those crashes, because a silent skip would shrink the fixture and change the split without telling the receipt.

So where did a model belong in this, if the checker is the whole grade? Not in the number. In the boring draft.

I got stuck on lowercase and empty strings while I was still half-listening to a lecture. I wanted a first pass at clean_row, and I did not want to spend the lab block installing a local stack just to strip a sentence. MonkeyCode is an open-source project, and the availability I am treating as operator-supplied is free model access plus a free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach.

I asked for a draft of the cleaner, then I pasted that draft beside mine and ran the empty-text case myself. I did not ask the model to choose the holdout. I did not ask it for an accuracy number, and I would not have kept one if it had offered.

The free server option mattered for a dull reason. This script is tiny, but my laptop was already warm from a video call, and I wanted the receipt somewhere that was not my battery. I am not going to invent a machine size, a token allotment, or a time limit in this note. Those numbers move, and a student blog is a terrible place to freeze them.

If you try the same path, read the project docs the same day you click. If free does not mean what you need this week, do not pretend a paragraph of mine updated the terms. A spare desk is useful. A stale promise is not.

What did the case show? The grouped receipt passed. The naive cut failed on s1. The claimed JSON failed on s2, and the blank text failed closed.

That is the result. No chart, and no leaderboard. A student lab that can say this split does not share ids, without borrowing confidence from a paragraph.

After you run both exits, you should be able to say what an id-level holdout protects, what a row cut leaks, and why a digest is only a lock on the ids you actually wrote down. If you cannot rebuild those ids, the metric is a rumor. Did the draft ever show you the ids, or only the adjective?

What I still do not trust is longer than what I trust. The receipt does not know that two different ids wrote the same sentence. It does not know that s1 and s2 are lab partners who shared a screenshot. It will happily put a later week in train and an earlier week in test if you only hand it ids.

The tail-slice rule is deterministic, which I want for a demo, and it is not a random stratified split, which a methods section might need. Hashing the id lists will not catch a flipped label inside a row. If your fixture holds real student writing, do not upload it to a hosted box just because the box is free.

Synthetic lines like mine are the whole dataset on purpose. Who should skip this approach? If you are writing a paper, this script is a preflight, not a split you can cite.

If the unit of leakage is a household, a course section, or a device, grouping only on student_id will bless a leak. If you need a GPU job, a serving stack, or a quota you can budget against, I did not test those, and I will not decorate this note with them. If a draft README wants a metric and you do not have a receipt yet, you are the person this is for.

The mistake I keep seeing in my own notes is treating a clean diff as a clean experiment. The other one is asking a model to make a proper train-test split and then grading the prose. A third, quieter one sits in the shuffle: you split, you shuffle, you save only the scores. Can you rebuild the ids from the score file?

If not, you do not have a lab. You have a mood. I have shipped moods before, and they do not survive the second reader.

If you want a next inch, add a second key. Group by student_id and section together, and make a fixture where the same student shows up under two sections. Does your receipt still pass? I would bet it passes for the wrong reason.

Send the smallest CSV that fools it. I am leaving cleaner drafts on the free-model side of the workflow, and the pass or fail on this script. If a free server is useful to you, use it as a spare desk, not as a witness.

The witness is the overlap list. Would you publish a number you cannot rebuild from that list?

Top comments (0)