Freeze current report bytes before you extract a parser. The smallest safe change keeps every known bug. A later commit can fix that known bug.
A messy function often writes files and parses flags together. Splitting it first can hide a behavior change. Characterization tests record the old contract before any split.
Why this order matters
Legacy helpers mix parsing, clocks, and disk writes. One rename can change error text by accident. Tests that only check happy paths miss that drift.
Michael Feathers named these behavior locks characterization tests. They lock observed behavior, including known product defects. They do not define the ideal future design.
The term comes from Working Effectively with Legacy Code. That book treats pins as discovery, not as specs. Use the pins to detect drift, not to praise the mess.
The sample mess
The sample below is a proposal, not a production module. It was not executed while drafting this article. Treat the outputs as the contract you must pin.
# Proposal only: this module was not executed in this draft.
import json
from pathlib import Path
def run_job(argv, out_dir):
if not argv or argv[0] != "emit":
return 2, "usage: emit <name>\n"
name = argv[1] if len(argv) > 1 else ""
if name.strip() == "":
return 3, "name missing\n"
path = Path(out_dir) / f"{name}.json"
payload = {"name": name, "ok": True}
text = json.dumps(payload, separators=(",", ":"))
path.write_text(text, encoding="utf-8")
return 0, f"wrote {path.name}\n"
Three observable facts matter before any code split. Exit codes must stay two, three, and zero. The stderr-style text must keep its trailing newline.
File bytes stay compact JSON with no extra space. Spaces inside the name stay inside the file name. Those quirks are recorded data, not cleanup tasks.
Step 1 — Record seams before you touch code
List every input that the function reads today. List every output that a caller can observe. Do not list the design you wish you had.
| Seam | Current observation | Safe to change now |
|---|---|---|
| argv missing or not emit | code 2 and usage text | no |
| blank or missing name | code 3 and name missing | no |
| name with inner spaces | file name keeps spaces | no |
| JSON body | compact separators, ok true | no |
| clock, random, network | unused in this function | n/a |
This seam table is the entire change budget. Anything unmarked in that table stays fully frozen. A parser extract may not edit those rows.
Step 2 — Add characterization tests
Use pytest and tmp_path for the file seam. Assert return codes, text, and exact file bytes. Do not normalize whitespace in the golden text.
Install pytest in the same interpreter you will pin. The tmp_path fixture is documented in the pytest how-to pages. Link those docs beside the new test file.
https://docs.pytest.org/en/stable/
# Proposal only: these tests were not executed in this draft.
from report_job import run_job
def test_usage_when_command_missing(tmp_path):
code, text = run_job([], tmp_path)
assert code == 2
assert text == "usage: emit <name>\n"
def test_missing_name_keeps_code_three(tmp_path):
code, text = run_job(["emit", " "], tmp_path)
assert code == 3
assert text == "name missing\n"
assert list(tmp_path.iterdir()) == []
def test_spaces_survive_in_file_name(tmp_path):
code, text = run_job(["emit", "Ada Lovelace"], tmp_path)
assert code == 0
path = tmp_path / "Ada Lovelace.json"
assert path.read_bytes() == b'{"name":"Ada Lovelace","ok":true}'
assert text == "wrote Ada Lovelace.json\n"
Run the full suite before the parser extract. A red test means the pin is wrong. Fix the failing pin, not the production code.
python -m pytest test_report_job.py -q
You should commit only the green characterization pins. That green commit becomes the safe refactor baseline. The next commit may move production code only.
Step 3 — Make the smallest safe change
Extract one pure parser and leave writes behind. Keep run_job as the only disk writer here. Call the new parser from that old function.
# Proposal only: keep writes inside the old function.
def parse_emit(argv):
if not argv or argv[0] != "emit":
return ("usage", None)
name = argv[1] if len(argv) > 1 else ""
if name.strip() == "":
return ("missing", None)
return ("ok", name)
def run_job(argv, out_dir):
kind, name = parse_emit(argv)
if kind == "usage":
return 2, "usage: emit <name>\n"
if kind == "missing":
return 3, "name missing\n"
path = Path(out_dir) / f"{name}.json"
payload = {"name": name, "ok": True}
text = json.dumps(payload, separators=(",", ":"))
path.write_text(text, encoding="utf-8")
return 0, f"wrote {path.name}\n"
Re-run the same three characterization tests without edits. Those tests must stay green without assertion edits. A needed assertion change means the extract was not safe.
python -m pytest test_report_job.py -q
git diff --stat
Review the diff for one function move only. Do not edit user-facing messages in this diff. Do not rename JSON keys in this same diff.
Do not add new flags in this same commit. Do not reorder imports unless a test requires it. Stop when the diff is only the parser move.
Use this gate before you accept the extract. Every row must pass in the same commit. A single miss means you revert the move.
| Check | Pass rule | Fail action |
|---|---|---|
| Three pins green | no assertion edits | revert the extract |
| Diff shape | one new function only | split the commit |
| Messages and bytes | identical to the pins | restore old text |
| Generated patch | tests only, else drop it | discard the patch |
Step 4 — Use a free model only as a draft aid
Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode offers free model access and a free server option. Those two facts are the only product claims used here.
Quotas, model names, and hardware are not assumed. Duration and package images are also not assumed. Check current terms before you rely on either option.
A free model may draft the three tests from the table. Use the free server option only if it can run your suite. You still compare its output with the seam table.
Give the model the current function and the seam table. Ask for tests that fail on any text change. Reject any patch that cleans the usage string text.
Draft characterization tests only from the seam table.
Do not change production code in this patch.
Assert exact codes, text, and file bytes.
Do not fix the space-in-name quirk at all.
Run that generated draft where your runner actually exists. Keep the command log next to the tests. If the server environment differs, record the Python version in notes.
python -V
python -m pytest test_report_job.py -q --tb=short
The draft model does not approve the merge. A human checks the diff against the table. Green tests plus a small diff are the gate.
Step 5 — Stop at the first safe boundary
Do not fix the space-in-name quirk in this pass. That space-in-name quirk stays pinned on purpose here. A behavior fix needs a new test and a new commit.
Do not extract the JSON writer in the same change. Two extracts hide which move broke a byte. One function move is enough for this pass.
Keep suggested commit subjects deliberately narrow and specific. Put the pin commit before the extract commit. Do not squash them until review is done.
test: pin emit codes, text, and report bytes
refactor: extract parse_emit without behavior change
What a bad extract looks like
A bad extract changes text while moving the parser. The usage line might lose its trailing newline. The missing-name code might become a raised exception.
Those edits can look cleaner during code review. The characterization suite should turn red immediately after. Revert the production edit and keep the pins.
Compare the red assertion to the seam table. The table row is the source of truth here. Do not update the pin to match the cleaner text.
Limits of this workflow
Characterization tests lock accidents as well as intent. A later reader may treat a bug as a feature. Name the quirk in the test docstring or a nearby comment.
Exact byte asserts break on harmless JSON key order changes. This sample uses one dict and explicit separators. This sample assumes insertion-ordered dicts from Python 3.7.
If key order is not a contract, compare a parsed object instead. Keep byte asserts only for text you truly freeze. Say that comparison choice in the test name.
The tmp_path fixture does not catch cross-process file locks. It also does not catch timezone-based file writes. Add those seams only if the function uses them.
A free server may not match production packages. Pin the test interpreter in the notes you keep. Do not treat a green remote run as a release sign-off.
Free model drafts can invent incorrect exact assertions. They can also strip important trailing newline bytes. Diff every generated line against the seam table.
Who should skip this approach
Skip it when the ticket requires a behavior change now. A pin would only slow the required fix. Write the new contract test first in that case.
Skip it when the repo has no isolated runner. Shared staging data can leak into golden pins. Build a local fixture path before you record goldens.
Skip it when outputs include secrets or personal data. Golden files would copy those values into git. Redact or synthesize inputs before you assert bytes.
Skip it when timing, network, or hardware is the contract. A free server is the wrong oracle for those seams. Measure those seams on the real target instead.
What to do next
Pick one messy function with three observable outputs. Pin those outputs, then extract one pure helper. Leave the bug fix for the following commit.
If you use a hosted assistant, read the current free-model and free-server terms first. Use them to draft and run pins, not to skip review.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.