A wide model diff is not a finished refactor. Lock observable behavior before you accept any split. Then ship only the smallest change that still matches.
Messy modules hide three jobs in one file. They read input, compute a score, and write logs. One careless edit can break all three jobs together.
This walkthrough uses a tiny Python scoring module. Every listing below is an unexecuted example for local use. Adapt names and paths before you run a command.
What must stay fixed
Readers care about outputs, not internal file shape. A characterization test records those outputs before the edit. The same test must pass after the smallest extract.
Record the return value, the log text, and any other bytes. Also record the exception type on the bad path. Skip fields you cannot reproduce on a second run.
Do not assert private helper names in this pass. Do not assert source line counts inside the module. Those checks punish a safe extract that preserves behavior.
Why a wide diff fails review
A coding model often proposes three moves together. It extracts helpers, renames files, and reorders imports. You cannot tell which move changed the score.
Treat review as a gate with a written rule. Reject the patch when it touches more than one concern. Ask for a new patch that moves one function only.
A concern here means load, score, log, or imports. Count them in the diff before you discuss style. Style nits can wait until the golden test is green.
Workflow you can repeat
Follow these steps in order on a clean branch. Stop when a step fails and do not skip ahead. Each step has one job and one pass condition.
1. Freeze the fixture set
Pick three inputs that already exist in the repo. Use one happy path, one edge path, and one bad path. Store them under fixtures with names that will not drift.
Do not invent production totals for this exercise. Do not copy customer files into the sample tree. Synthetic rows are enough to lock the current contract.
2. Write the characterization test
Call the public function and capture every side effect. Compare the capture with a golden file you just recorded. Fail the test when any captured field drifts.
Keep the golden file in version control with the test. Reviewers should see the expected bytes beside the patch. A missing golden file means the gate is not ready.
3. Run the test on unchanged code
Confirm the golden file matches the current module. Treat a red test here as a fixture bug, not a product bug. Fix the fixture before you touch production code.
Run the test twice and require the same result. A flaky capture is not a valid characterization ledger. Remove clocks and random calls from this fixture first.
4. Reject the wide model diff
Read the diff and count the concerns it changes. If the diff changes two concerns, discard that patch. Request one extract of the pure score function only.
Do not negotiate a partial apply of a wide diff. Partial applies hide the hunk that actually broke output. The model can draft again against a tighter prompt.
5. Apply the smallest safe change
Move the pure score function into a new module. Leave file reads and log writes in the old module. Keep the public function name and the argument order.
Update only the import that the old module needs. Do not reformat unrelated functions in the same commit. Formatting noise makes the next review slower and weaker.
6. Re-run the same characterization test
Do not edit the golden file after the extract. A green result means observable behavior stayed fixed. A red result means you revert the extract immediately.
Inspect the assertion diff before you edit production code again. If the log text changed, the extract was not pure. Put the function back and repeat the prompt with a harder limit.
7. Consider a second extract only after green
Repeat the gate for logging or file reads later. Never stack a second extract on a red ledger. Each pass should change one concern and nothing else.
Wait for a human review between those passes. A model can suggest the next extract, not approve it. Approval stays with the person who owns the module.
Example module before the split
This listing is a proposal, not a measured benchmark. It shows the tangle you should not rewrite in one pass. No runtime numbers are claimed for this sample.
# Unexecuted example. Adapt paths before any local run.
import json
from pathlib import Path
def load_rows(path):
raw = Path(path).read_text(encoding="utf-8")
return json.loads(raw)
def score_rows(rows):
total = 0
for row in rows:
total += int(row["points"]) * int(row["weight"])
return total
def write_log(path, message):
Path(path).write_text(message + "\n", encoding="utf-8")
def run_job(src, log_path):
rows = load_rows(src)
total = score_rows(rows)
write_log(log_path, f"total={total}")
return total
The public job mixes load, score, and log writes. The pure function is score_rows and nothing else. That function is the only safe first extract.
Integer casts can raise ValueError on bad rows. Capture that exception type in a separate bad-path test. Do not swallow the error while you move the function.
Characterization test you can copy
This test is also an unexecuted local example. It uses the standard library and makes no network calls. Run it with pytest after you place both files.
# Unexecuted example. Requires pytest in your local environment.
import json
from pathlib import Path
import score_job
def test_happy_path_matches_golden(tmp_path):
src = tmp_path / "rows.json"
log_path = tmp_path / "job.log"
src.write_text(
json.dumps([{"points": 2, "weight": 3}]),
encoding="utf-8",
)
result = score_job.run_job(src, log_path)
golden = {"result": 6, "log": "total=6\n"}
actual = {
"result": result,
"log": log_path.read_text(encoding="utf-8"),
}
assert actual == golden
def test_empty_rows_score_zero(tmp_path):
src = tmp_path / "empty.json"
log_path = tmp_path / "job.log"
src.write_text("[]", encoding="utf-8")
result = score_job.run_job(src, log_path)
assert result == 0
assert log_path.read_text(encoding="utf-8") == "total=0\n"
def test_bad_row_raises_value_error(tmp_path):
src = tmp_path / "bad.json"
log_path = tmp_path / "job.log"
src.write_text(
json.dumps([{"points": "x", "weight": 1}]),
encoding="utf-8",
)
try:
score_job.run_job(src, log_path)
except ValueError:
return
raise AssertionError("expected ValueError")
The happy-path total is 2 times 3, which is 6. The empty input total is 0 by the same loop. Those two numbers are arithmetic checks, not benchmarks.
The happy-path assertion checks the result and the log text. It does not check helper names or file layout. A later extract can move score_rows and still pass.
The bad-path test locks the exception type only. It does not lock the exception message text. Message text often changes during a harmless rename.
Decision table for the incoming patch
Use this table before you apply any model patch. Count concerns in the diff, then pick one row. The table is a review rule, not a product score.
| Diff shape | Concerns touched | Decision |
|---|---|---|
| Extract score_rows only | 1 | Apply, then re-run the golden tests |
| Extract plus a rename | 2 | Reject and request a smaller patch |
| Extract plus import reorder | 2 | Reject until imports are a later pass |
| New dependency added | 2 or more | Reject for this characterization pass |
| Golden file edited by the model | n/a | Reject and restore the recorded golden |
| Behavior change requested in the prompt | n/a | Stop and write a new spec instead |
A one-concern row is the only apply decision. Every other row waits for a narrower diff. Restore the branch if a rejected hunk was already applied.
Where a narrow model draft fits
Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode, as supplied for this draft, offers free model access. That supplied note also includes a free server option.
Ask the model only to draft the one-function extract. Use the free server only when your current terms allow it. You still own the accept or reject decision.
Treat the prompt below as a proposal, not a guarantee. It limits the model to one concern and one new file. Edit the names so they match your module.
Extract only score_rows into score_math.py.
Do not rename run_job.
Do not reorder imports in other files.
Do not edit tests or golden files.
Return a diff that changes one concern.
Paste the model diff into your review branch, not into main. Run the characterization tests before you create any commit. Discard the diff when the table says reject.
If you already have MonkeyCode access, try this gate on one messy module. Keep the patch only when both golden tests stay green. Leave the wide rewrite for a later, separate pass.
Commands that keep the loop small
These commands assume a local checkout and a pytest install. They are examples, not proof of a remote execution. Replace the test path if your layout differs.
python -m pytest tests/test_score_job.py -q
git diff --stat
git diff --name-only
Read the name list before you read the hunks. One new module plus one caller edit is the target shape. More files mean you stop and request a smaller patch.
A quiet pytest run is the pass signal for this gate. A failed assertion is a stop signal, not a prompt to improvise. Revert first, then ask the model for a narrower diff.
Limitations you should state in review
Characterization tests miss behavior that you never recorded. A green run does not prove thread safety or speed. It only proves the captured fields stayed the same.
This sample ignores permissions, clocks, and network calls. Add those captures before you extract code that uses them. Do not treat this fixture as a full production suite.
Free model access can still propose a wide diff. A free server can still execute the wrong test file. Neither option replaces the concern count in the table.
This draft states no speed, cost, or uptime figures. Those facts were not supplied and must not be guessed. Check current product terms yourself before you depend on access.
The integer score rule is a teaching stand-in, not your domain rule. If your formula uses floats, lock rounding in the golden file. Otherwise a harmless refactor can look like a behavior change.
Who should skip this approach
Skip this gate when the module has no stable outputs. Skip it when the next change must alter behavior on purpose. Write a new spec first if the score rule itself must change.
Skip it for security fixes that need a behavior change. A characterization test would freeze the bug you must remove. Use a failing regression test for that case instead.
Skip it when you cannot run tests on the target branch. A free server does not help if secrets block the job. Do not paste private fixtures into any hosted prompt.
Skip it when the team has no reviewer for the final diff. A green test is necessary, but it is not the whole review. Ownership of the merge stays with a person, not a model.
Closing rule
Start from the conclusion and keep every pass small. Lock outputs, reject wide diffs, and extract one function. Commit only after the same golden tests stay green.
Top comments (0)