A green test run can still be a failed repair, and that is the failure this notebook is meant to catch. The patch looks small, the command exits clean, and the assertion now expects the broken behavior. Have you shipped a change like that because the suite finally stopped shouting? I would not trust the next model-written diff until it survived the forty-eight hour routine below.
What I am trying to catch
I am not trying to prove that a model is wise, and I am not trying to ban model patches either. I am trying to see whether the diff repaired production code or quietly edited the oracle that judges it. An oracle here means the assertion, the snapshot, the golden file, or the fixture the test compares against. If that oracle moves, a green run can mean the test learned to expect the bug.
Disclosure: This article was prepared as part of MonkeyCode's product outreach. If a free model is already available on your account, I would use it only to draft a unified diff, and I would spend the free server option only on the single clean command later in this notebook. I am not naming a model, a quota, a machine size, or a promise that either option stays available.
The three lines I freeze first
Before I ask for a patch, I write three lines in the notebook and I do not let the prompt negotiate them. The production paths that may change are listed first, and the test paths that must stay still are listed second. The third line is the exact command that failed, copied from the terminal rather than paraphrased from memory. Would you trust a summary of a failing command more than the command itself?
- Allow edits only under the production package path you name in the note.
- Deny edits under tests, snapshots, goldens, cassettes, and expected-output files.
- Save the model reply as a
.difffile and do not apply it yet. - Classify that diff locally, then decide whether any test run is even worth starting.
A full-file rewrite is already a miss, even when the new file would pass. The noise hides the oracle edit, and I would ask again for a unified diff before I read any further. I keep the rejected reply in the same folder so the second draft has something honest to be compared against.
An unexecuted classifier you can read in one sitting
The script below is a starter, not a log from a finished incident, and it never calls a network. It uses only the standard library so you can audit every branch on a laptop with no extra install. It answers a narrow question: did the hunks land on production files, test files, or oracle-shaped paths, and did a test file gain an assertion line?
"""Unexecuted starter. Classify a unified diff before applying a model patch."""
from __future__ import annotations
import re
import sys
from dataclasses import dataclass, field
TEST_HINTS = ("test_", "_test.", "/tests/", "/test/", "spec/")
ORACLE_HINTS = (".snap", ".golden", "expected", "fixture", "cassettes/")
ASSERT_RE = re.compile(r"^\+\s*(assert |self\.assert|expect\(|pytest\.raises)")
@dataclass
class DiffReport:
prod_hunks: int = 0
test_hunks: int = 0
oracle_hunks: int = 0
added_assertions: int = 0
files: list[str] = field(default_factory=list)
def classify_diff(diff_text: str) -> DiffReport:
report = DiffReport()
is_test = False
is_oracle = False
for line in diff_text.splitlines():
if line.startswith("diff --git "):
current = line.split()[-1]
current = current[2:] if current.startswith("b/") else current
report.files.append(current)
low = current.lower()
is_test = any(hint in low for hint in TEST_HINTS)
is_oracle = any(hint in low for hint in ORACLE_HINTS)
continue
if line.startswith("@@"):
if is_oracle:
report.oracle_hunks += 1
elif is_test:
report.test_hunks += 1
else:
report.prod_hunks += 1
continue
if is_test and ASSERT_RE.match(line):
report.added_assertions += 1
return report
if __name__ == "__main__":
text = open(sys.argv[1], encoding="utf-8").read()
print(classify_diff(text))
I would keep that file beside the notebook, not inside the product package, so a later patch cannot quietly "fix" the judge. The counters are hints, not a proof, and a rename that omits diff --git will fall through as if nothing happened. That miss is why the hand-read row exists in the decision table further down.
A constructed diff, not a production incident
I do not have a measured corpus, a customer log, or a personal failure count to attach to this sample. The diff below is constructed so you can see the shape of a bad patch before you point the script at your own tree. Replace it with a diff from your saved failing command as soon as the script runs locally.
diff --git a/pkg/parse.py b/pkg/parse.py
--- a/pkg/parse.py
+++ b/pkg/parse.py
@@ -10,1 +10,1 @@
- return raw
+ return raw.strip()
diff --git a/tests/test_parse.py b/tests/test_parse.py
--- a/tests/test_parse.py
+++ b/tests/test_parse.py
@@ -4,1 +4,1 @@
- assert parse(" x ") == " x "
+ assert parse(" x ") == "x"
On this sample, prod_hunks and test_hunks should both be at least one, and added_assertions should be at least one. oracle_hunks stays zero because the test path does not match the golden-file hints, which is the point of keeping a separate assertion counter. A local run would go green after you applied both sides, and that green would forget the old contract. Can you explain the new expected value without pointing at the patch itself?
python classify_diff.py samples/oracle_edit.diff
If you want a tiny self-check, save the sample and assert the four counters in a separate file that the model is not allowed to edit. I would fail the check when added_assertions is zero on this sample, because that would mean the regular expression drifted. I would not treat a passing self-check as evidence about any real repository.
How I would spend the forty-eight hours
This schedule is a proposal for the next time a model offers a patch, not a diary of a run I already finished. I would repeat the order, and I would not repeat any test command that skipped the classifier. The clock is a cap so the notebook ends, not a claim that the work fills every hour.
Hours 0 to 8: freeze the contract
I capture one failing command, one allow-list, and one deny-list, and I store them in the same note as the date. I ask for a unified diff only, and I paste the deny-list in plain language beside the traceback. If the reply rewrites a whole file, I stop and ask again, because a full-file rewrite hides the oracle edit in unrelated noise. I do not open the editor on production code during this block.
Hours 8 to 24: classify before you execute
I run the classifier on every saved diff and I write the four counters next to the filename in the note. Only a diff with production hunks, zero oracle hunks, and zero added assertions earns a local test command. A diff that touches tests gets a second prompt that calls the test path read-only, and I keep both replies. Would a second draft that still edits the assertion change your mind, or only your patience?
python classify_diff.py patch.diff
# Only if oracle_hunks is 0, added_assertions is 0, and files match the allow-list:
python -m pytest tests/test_parse.py -q --tb=short
That pytest line is an illustration, and your notebook should store the command that actually failed. If the project uses unittest or a custom runner, write that runner down instead of borrowing mine. I would also record the working directory, because a green run from the wrong folder is a different bug.
Hours 24 to 48: one clean run, then stop
Some failures are local accidents: a developer install, a shell alias, or a file that exists only on your laptop. That is the only moment I would spend a free server, and I would spend it on the same command I already wrote down. I would not retune the prompt on the server, and I would not treat one green remote log as a new contract. I copy the remote log back into the note, then I return to the classifier before anyone talks about merging.
A decision table, not a vibe
| What the classifier shows | What I do next | What I refuse to do |
|---|---|---|
| Production hunks only, command is local | Run the saved test command on my machine | Ask for a broader rewrite |
| Test hunks, oracle hunks, or new assertions | Reject the diff and re-prompt with a read-only test path | Celebrate a green suite |
| Failure needs a clean machine, diff is clean | One free-server run of the saved command | Debug by editing the server image |
| Diff is not unified, or files are renamed | Read the patch by hand and ignore the counters | Trust a zero as proof of safety |
What breaks if you trust the hints too far
The hints are strings, so a production path that contains test will look guilty, and a test named check_parse.py will look innocent. Assertion helpers with your own names will not match the regular expression, and a binary fixture will not show up as a text hunk at all. A clean server can hide a local-only bug, or invent a new one, when its image is not your image. I would not use path hints alone on a tree that generates tests into the production package.
There is a second failure mode I would write in the margin before I close the notebook. A model can leave the test file untouched and still weaken the contract by catching a broader exception in production code. The classifier will call that a production hunk, which is correct and also incomplete. That is why a human still has to read the allowed diff, even after every counter looks boring.
Who should skip this approach
I would not use this notebook for a security patch, a data migration, or any change where a missed oracle edit has legal or safety weight. Teams that need a formal proof, or a review gate with an explicit file allow-list from the build system, should use that gate instead of these string hints. People who do not have a failing command they can re-run should not start the clock at all. A free server is not a durable CI replacement in this note, because I am not claiming how long it exists or what image it boots.
What I would repeat next week
I would repeat the deny-list, the saved command, and the habit of classifying a diff before applying it. I would not repeat a server run that I cannot describe in one sentence, and I would not repeat a prompt that forgot to forbid test edits. If the second draft still touches the oracle, the bug is in my contract, not in the button I pressed. The useful close is boring on purpose: keep the patch, keep the counters, and let a person accept a new expected value only when they can say why the old one was wrong.
Top comments (0)