A scored cut beats a wide cleanup every time. Pick the change with the lowest caller risk. Golden tests then guard that one extract, and nothing else.
This sample is a messy shift-summary script for practice. The listings are proposed examples and were not executed here. They are not benchmarks, customer cases, or measured timings.
Why the wide edit loses
Stdout, files, and return values move together in a broad patch. You cannot name the line that changed the total. A scored cut keeps two of those surfaces still.
Private helpers can wait through this first pass. Callers see printed lines, return fields, and CSV rows. Those three fields are the contract you pin.
The messy module
Clock reads, CSV writes, and row cleanup share one function. That mix is why a rewrite feels tempting. You will not split all three in this pass.
# proposed example — unexecuted; adapt names before use
import csv
from datetime import datetime, timezone
from pathlib import Path
def build_shift_report(rows, dest):
dest = Path(dest)
dest.parent.mkdir(parents=True, exist_ok=True)
stamped = []
for row in rows:
name = str(row.get("name", "")).strip()
hours = row.get("hours", 0)
try:
hours = float(hours)
except (TypeError, ValueError):
hours = 0.0
if hours < 0:
hours = 0.0
stamped.append((name, f"{hours:.2f}"))
with dest.open("w", newline="", encoding="utf-8") as handle:
writer = csv.writer(handle)
writer.writerow(["name", "hours", "built_at"])
built = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
for name, hours in stamped:
writer.writerow([name, hours, built])
total = sum(float(hours) for _, hours in stamped)
print(f"rows={len(stamped)} total={total:.2f} path={dest}")
return {"rows": len(stamped), "total": round(total, 2), "path": str(dest)}
The timestamp changes on every plain unpinned run. A raw file compare would fail without a frozen clock. Patch the clock before you store an expected row.
Step 1. Score cuts before editing
List candidate edits in a table before you touch code. Score caller-visible risk, not how elegant the helper looks. Keep only the lowest-risk pure cut in this pass.
| Candidate change | Caller-visible risk | Decision |
|---|---|---|
| Extract row cleanup into format_row | Low if the writer stays | Keep |
| Replace csv.writer with manual joins | High because quoting can change | Drop |
| Rename headers or the dest argument | High for downstream readers | Drop |
| Drop the zero floor on bad hours | High because totals change | Drop |
| Fold the print into the return dict | Medium because stdout changes | Drop |
The keep row is the smallest safe change. Paths, headers, and the zero floor stay put. A later pass can score those rows again.
Step 2. Pin the current contract
Write one characterization test against two fixed rows. Freeze datetime.now at the import path the module uses. Compare parsed CSV rows, not a single raw blob.
Parsed rows avoid a false failure from line endings. Python's csv module defaults to carriage-return line endings. A byte-identical lock needs an explicit lineterminator, which this pass does not add.
# proposed test — unexecuted; check fixture names in your pytest
import csv
from datetime import datetime, timezone
from shift_report import build_shift_report
def test_shift_report_golden(tmp_path, capsys, monkeypatch):
class Frozen(datetime):
@classmethod
def now(cls, tz=None):
return cls(2024, 1, 2, 3, 4, 5, tzinfo=timezone.utc)
monkeypatch.setattr("shift_report.datetime", Frozen)
dest = tmp_path / "out" / "shift.csv"
rows = [
{"name": " ada ", "hours": "1.5"},
{"name": "bo", "hours": "nope"},
]
result = build_shift_report(rows, dest)
printed = capsys.readouterr().out.splitlines()
assert result == {"rows": 2, "total": 1.5, "path": str(dest)}
assert printed == [f"rows=2 total=1.50 path={dest}"]
with dest.open(newline="", encoding="utf-8") as handle:
parsed = list(csv.reader(handle))
assert parsed == [
["name", "hours", "built_at"],
["ada", "1.50", "2024-01-02T03:04:05Z"],
["bo", "0.00", "2024-01-02T03:04:05Z"],
]
This test records today's behavior, including the zero floor. It is not approval of that payroll rule. Update it only when a product owner wants a new contract.
Confirm tmp_path, capsys, and monkeypatch against your installed pytest. Those three names are long-standing public pytest fixtures. This article does not pin a release number.
Step 3. Run the pin twice
Run the golden test and save the first result text. Fix the harness until the same command passes on unchanged code. Do not start the extract while the pin is red.
python -m pytest tests/test_shift_report.py::test_shift_report_golden -q
Run that same command once more before you edit. A matching green result is your only baseline. A changed failure means the harness is still unstable.
Step 4. Apply the keep row only
Move name cleanup and hours flooring into format_row. Leave mkdir, csv.writer, print, and datetime.now in the original function. Pass no clock value into the pure helper.
# proposed extract — apply only after the golden test is green
def format_row(row):
name = str(row.get("name", "")).strip()
hours = row.get("hours", 0)
try:
hours = float(hours)
except (TypeError, ValueError):
hours = 0.0
if hours < 0:
hours = 0.0
return name, f"{hours:.2f}"
Call format_row from inside the existing row loop. Keep the stamped list and the writer block as they are. If the diff touches dest, headers, or print, revert it.
Add unit cases for the new helper after the extract. Cover a blank name, a negative number, and a bad string. Keep the golden test in the same file.
# proposed unit checks — unexecuted; add this import after the extract
from shift_report import format_row
def test_format_row_edges():
assert format_row({"name": " ", "hours": -2}) == ("", "0.00")
assert format_row({"name": "cy", "hours": "bad"}) == ("cy", "0.00")
assert format_row({"hours": "2"}) == ("", "2.00")
Step 5. Draft on a free model, judge locally
A free coding model may draft format_row from the table and the golden test. The draft is a proposal, not an approval. You reject any edit outside that helper and its call site.
Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode's free model access can supply that draft. Its free server option can run this pytest file away from production paths.
Those two notes are operator-supplied availability claims, not quotas, hardware specs, durations, or permanence. Send the model the golden test, the cut table, and the current function. Ask for a patch that leaves the three asserts unchanged.
python -m pytest tests/test_shift_report.py -q
Run the suite yourself and accept only a green result. Do not paste secrets, live exports, or customer names into the prompt.
The two fixture rows are enough model context. A free server is still remote, so synthetic fixtures remain mandatory.
Step 6. Triage failures before a second edit
Use the same order each time a run goes red. Match the symptom, then apply one next action. Do not stack a second cleanup on a red pin.
| Symptom | Likely cause | Next action |
|---|---|---|
| Timestamp cell drifts | Clock patch missed the import | Patch shift_report.datetime and rerun |
| Total or hours text drifts | format_row changed the floor | Revert the helper and compare |
| Header or path drifts | Diff widened past the keep row | Discard the draft |
| Unit test fails, golden passes | New edge is outside old rows | Fix the helper or the new case only |
If the golden parsed rows change, stop immediately. Restore the writer and the zero floor next. Do not refresh the expected rows to hide an unplanned edit.
Limits of this pass
Golden tests preserve behavior callers already depend on. The zero floor may be wrong for payroll math. Keep it until an owner accepts a new expected table.
Float sums are not a money type in this sample. Two rows hide rounding errors that more rows would show. Pin a decimal helper only in a later scored cut.
This pass does not prove thread safety, disk-full behavior, or encoding errors. Those need separate pins if your callers rely on them. One green file is not a full audit.
No runtime, token cost, or accuracy figure is reported here. No model name is stated, because none was verified for this draft. Confirm the current free-access offer with the vendor before a team depends on it.
The datetime subclass is a teaching seam for one module. Code that imports datetime in several files needs one patch per path. A missed path makes the timestamp cell flake.
Who should not use it
Skip this flow when this week's release must change the CSV contract. A golden lock would block the intended rows. Write the new expected table first, then code to it.
Skip it when the old function cannot run without production damage. Add a dry-run seam before any characterization pin. Never aim the first test at live folders.
Skip it when the goal is a framework swap in one branch. The table exists to refuse that wide cut. Split the swap into later scored rows, or schedule a rewrite.
Skip model drafting when you lack synthetic rows for regulated data. Remote runs do not make secret inputs safe. Local pytest with fake rows is the safer default.
Close
Score the cuts, pin parsed rows, and extract one pure helper. Re-run the same golden test before any second edit. A model may draft the patch, but the suite decides.
Point MonkeyCode's free server at this pytest file if you already have it. Keep the extract only when that run stays green.
Top comments (0)