DEV Community

Dakota Huang
Dakota Huang

Posted on

Golden-Master a Tangled Script Before the Cleanup Diff

Do not clean a tangled script before you freeze outputs. Write characterization tests against today's actual messy behavior. Only then apply the smallest safe change.

AI diffs make local edits cheap and frequent. They do not make observable behavior cheap to verify. A messy repo hides effects in prints, files, and globals.

This article is a labeled worked example, not production history. The protocol stays useful without any coding assistant. Cheap model output does not replace a checked-in oracle.

The failure mode

A typical messy script mixes calculation, I/O, and formatting. Helpers share one mutable dictionary across branches. Exit codes often depend on print order.

Tests are missing, or they mock every call. An assistant then rewrites the whole file. The diff looks small and deceptively tidy.

Downstream jobs then break on whitespace, paths, or codes. The missing artifact is a behavior oracle. The oracle is current output, not intended design.

Step 1: Inventory the tangle

Work on one entrypoint at a time. Do not start the cleanup inside helpers. Capture four facts before any source edit.

  1. Command that operators actually run today.
  2. Files the process reads and writes.
  3. Stdout, stderr, and numeric exit code.
  4. Environment variables that flip control flow.

Record the entrypoint in a checked-in text file. Keep that file next to the characterization tests. Do not trust chat memory for this inventory.

Proposed inventory commands for the sample script:

mkdir -p tests/golden
git rev-parse HEAD
git status --short
python3 -c "import sys; print(sys.version)"
Enter fullscreen mode Exit fullscreen mode

Proposed tests/inventory/score_jobs.txt contents:

entrypoint: python3 score_jobs.py jobs.csv
reads: jobs.csv
writes: $SCORE_OUT/summary.txt, $SCORE_OUT/failures.json
stdout: one status line per job
stderr: empty on the happy path
exit: 0 if any job scored, 2 if none scored
env: SCORE_STRICT=1 treats unknown status as failure
env: SCORE_OUT selects the output directory
Enter fullscreen mode Exit fullscreen mode

Step 2: Isolate a fixture

Copy a tiny input next to the test. Do not copy production exports into the fixture. Keep the fixture rows ugly on purpose.

Proposed fixtures/jobs.csv:

id,status,weight
a1,done,2
b2,UNKNOWN,1
c3,done,0
d4,failed,4
Enter fullscreen mode Exit fullscreen mode

Ugly rows are the contract, not noise. Zero weights and unknown statuses encode real branches. Clean sample data hides the mess you must pin.

Step 3: Record goldens once

Write a recorder before you write assertions. The recorder must run the real entrypoint. It dumps stdout, stderr, exit code, and files.

Proposed tools/record_score_jobs_oracle.py:

"""Record golden outputs for score_jobs.py. Proposed example only."""
from __future__ import annotations

import os
import subprocess
import sys
from pathlib import Path

ROOT = Path(__file__).resolve().parents[1]
SCRIPT = ROOT / "score_jobs.py"
FIXTURE = ROOT / "fixtures" / "jobs.csv"
GOLDEN = ROOT / "tests" / "golden" / "score_jobs"


def record(label: str, env_extra: dict[str, str]) -> None:
    tmp = ROOT / ".oracle-tmp" / label
    if tmp.exists():
        for child in tmp.rglob("*"):
            if child.is_file():
                child.unlink()
    out_dir = tmp / "out"
    out_dir.mkdir(parents=True, exist_ok=True)
    env = os.environ.copy()
    env["SCORE_OUT"] = str(out_dir)
    env.update(env_extra)
    proc = subprocess.run(
        [sys.executable, str(SCRIPT), str(FIXTURE)],
        cwd=tmp,
        env=env,
        text=True,
        capture_output=True,
        check=False,
    )
    dest = GOLDEN / label
    dest.mkdir(parents=True, exist_ok=True)
    (dest / "exit").write_text(str(proc.returncode))
    (dest / "stdout.txt").write_text(proc.stdout)
    (dest / "stderr.txt").write_text(proc.stderr)
    summary = out_dir / "summary.txt"
    failures = out_dir / "failures.json"
    (dest / "summary.txt").write_text(
        summary.read_text() if summary.exists() else ""
    )
    (dest / "failures.json").write_text(
        failures.read_text() if failures.exists() else ""
    )


if __name__ == "__main__":
    record("default", {})
    record("strict", {"SCORE_STRICT": "1"})
    print("wrote", GOLDEN)
Enter fullscreen mode Exit fullscreen mode

Run the recorder on an unchanged tree. Read every golden file by hand before commit. Commit those golden files as the behavior contract.

python3 tools/record_score_jobs_oracle.py
git add tests/golden/score_jobs fixtures/jobs.csv
git status --short
Enter fullscreen mode Exit fullscreen mode

Do not re-record after a cleanup diff. Re-recording after a cleanup hides real regressions. Update goldens only after a product decision.

Step 4: Assert the dirt

The test reuses the same runner shape. It compares exact bytes, not review vibes. It should fail on a single space.

Proposed tests/test_score_jobs_oracle.py:

"""Characterization tests for score_jobs.py. Proposed example only."""
from __future__ import annotations

import os
import subprocess
import sys
from pathlib import Path

import pytest

ROOT = Path(__file__).resolve().parents[1]
SCRIPT = ROOT / "score_jobs.py"
FIXTURE = ROOT / "fixtures" / "jobs.csv"
GOLDEN = ROOT / "tests" / "golden" / "score_jobs"


def run_script(tmp: Path, extra: dict[str, str] | None = None) -> dict:
    out_dir = tmp / "out"
    out_dir.mkdir()
    env = os.environ.copy()
    env["SCORE_OUT"] = str(out_dir)
    if extra:
        env.update(extra)
    proc = subprocess.run(
        [sys.executable, str(SCRIPT), str(FIXTURE)],
        cwd=tmp,
        env=env,
        text=True,
        capture_output=True,
        check=False,
    )
    summary = out_dir / "summary.txt"
    failures = out_dir / "failures.json"
    return {
        "exit": proc.returncode,
        "stdout": proc.stdout,
        "stderr": proc.stderr,
        "summary": summary.read_text() if summary.exists() else "",
        "failures": failures.read_text() if failures.exists() else "",
    }


@pytest.mark.parametrize(
    "label,extra",
    [("default", None), ("strict", {"SCORE_STRICT": "1"})],
)
def test_path_matches_golden(tmp_path: Path, label: str, extra: dict | None) -> None:
    actual = run_script(tmp_path, extra)
    expected = GOLDEN / label
    assert actual["exit"] == int((expected / "exit").read_text())
    assert actual["stdout"] == (expected / "stdout.txt").read_text()
    assert actual["stderr"] == (expected / "stderr.txt").read_text()
    assert actual["summary"] == (expected / "summary.txt").read_text()
    assert actual["failures"] == (expected / "failures.json").read_text()
Enter fullscreen mode Exit fullscreen mode

Run pytest on the harness before you touch score_jobs.py. A red harness means the recorder and test disagree. Fix that mismatch before any refactor work.

python3 -m pytest tests/test_score_jobs_oracle.py -q
Enter fullscreen mode Exit fullscreen mode

Step 5: Classify each edit

Do not accept a refactor plan as prose. Classify every intended change against the oracle.

Change idea Touches oracle? First commit? Next action
Rename a local variable No Yes Apply after tests pass
Extract a pure score helper No, if prints stay Yes Keep I/O in main
Reorder stdout lines Yes No Reject or retarget product
Change JSON indent or key order Yes No Freeze json.dumps as-is
Move file writes into a helper Maybe No Second commit, same goldens
Drop the unknown-status branch Yes No Needs an explicit spec test
Add type hints only No Yes Keep runtime identical

The first cleanup commit may only include off-oracle rows. If a row flips on, split the work. Do not bargain with the table in chat.

Step 6: Apply one safe line

Here is the proposed messy module for this walkthrough. Treat the file as unlabeled sample code.

# score_jobs.py — proposed messy entrypoint
import csv
import json
import os
import sys


def main():
    path = sys.argv[1]
    out = os.environ.get("SCORE_OUT", "out")
    strict = os.environ.get("SCORE_STRICT") == "1"
    os.makedirs(out, exist_ok=True)
    rows = list(csv.DictReader(open(path)))
    scores = []
    failures = []
    total = 0
    for row in rows:
        status = row["status"]
        weight = int(row["weight"])
        if status == "done":
            s = weight * 10
            scores.append((row["id"], s))
            total += s
            print("ok", row["id"], s)
        elif status == "failed":
            failures.append(row)
            print("fail", row["id"])
        else:
            if strict:
                failures.append(row)
                print("fail", row["id"], "unknown")
            else:
                print("skip", row["id"], status)
    open(os.path.join(out, "summary.txt"), "w").write(str(total) + "\n")
    open(os.path.join(out, "failures.json"), "w").write(json.dumps(failures))
    sys.exit(0 if scores else 2)


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

The smallest safe change extracts the arithmetic only. Leave prints and file writes inside main.

def score_done(weight: int) -> int:
    return weight * 10
Enter fullscreen mode Exit fullscreen mode

Replace s = weight * 10 with s = score_done(weight). Run the golden tests after that single replace. Stop if stdout, files, and exit still match.

Do not extract main in the same commit. Do not introduce a class for taste. Do not pretty-print the JSON payload yet. Those edits need their own rows in the table.

Step 7: What a model may draft

A model can draft the recorder from the inventory file. It can also propose the one-line extract. It cannot own the oracle or the table.

MonkeyCode offers free model access and a free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Paste the inventory, the fixture, and the messy file. Ask for a characterization harness, not a rewrite.

Keep these rules with any assistant:

  1. Generate tests first. Reject rewrite-only answers.
  2. Run the harness on your own machine.
  3. Diff golden files, not model prose.
  4. Allow one behavioral risk per commit.
  5. Drop the session if it edits goldens to match a new diff.

The free server is optional for this protocol. A local pytest run remains the source of truth. If the model output disagrees with goldens, keep the goldens.

Limitations

Golden masters pin bugs as well as features. That is the method, not a defect. They fail on timestamps, random IDs, and unordered sets.

This method does not prove functional correctness by itself. It only proves stability of current edges. Product intent still needs explicit tests later.

Large binary outputs do not belong in git goldens. Store a sha256 of those files instead. Live network calls need a recorded fixture. Never hit the live network during these tests.

Line endings and locale will break naive string compares. Normalize newlines in the recorder if your team mixes OS images. Do not normalize away spaces that operators already depend on.

Who should skip this

Skip this protocol for greenfield code with a written spec. Skip it when current behavior is unsafe or destructive. Skip it when snapshots would store secrets.

Do not use a model to fix failing goldens. That choice hides the regression you needed to see. Do not batch five extracts into one assistant diff. The table exists to stop that collapse.

After the first green commit

Tag the commit as a characterization baseline. Keep the inventory file on that commit. The next extract starts from the same goldens.

If later work must change stdout, add a spec test first. Then update goldens in a dedicated commit. Never mix format changes with logic changes.

Cheap generation does not retire this test sequence. Cheap edits make the sequence more necessary. Messy repos fail at the observable edges. Pin those edges, then cut one line.

Top comments (0)