A weekend contributor clones a small library, reproduces a parser bug, and asks a coding model to repair it before dinner. The model returns a tidy diff, a confident summary, and a test that already passes on the patched tree. The maintainer later discovers that the same test also passed on the unpatched tag, so the pull request never proved a fix. That pattern wastes review time, and a frozen oracle stops it before the branch leaves the fork.
Why a passing test is not evidence
An open-source review needs a failure that exists before the patch and disappears after the patch, under the same command. A generated test written after the edit can encode the new behavior without ever having failed. Reviewers then argue about intended behavior instead of comparing two captured runs from the same command. The workflow below treats those two runs as the contract that every later comment must respect.
Run the loop in order
- Write the manifest with an argument list, both expected statuses, and two markers before any production edit.
- Capture the before side on the untouched checkout and stop if the status is already the success code.
- Apply the smallest patch that can clear the failure marker, whether a person or a model drafted the diff.
- Capture the after side, run the checker, and refuse the pull request text when the checker returns an error.
- Ask the model to compare the diff with those files, and leave the merge decision with a human reviewer.
Lay out the evidence before editing
The contributor creates an evidence directory before touching production code, then records the failing command in a small manifest. That manifest names the command, the expected nonzero status, and a short marker that must appear in the failure output. A later checker refuses to bless a patch when either capture is missing or when the command string has changed. The point is narrow: prove this bug, on this command, before any model rewrites the story.
Directory contract
evidence/
manifest.json
before/stdout.txt
before/stderr.txt
before/status.txt
after/stdout.txt
after/stderr.txt
after/status.txt
Proposed manifest
The sample below is a proposal for one pytest node, and it is not a result from a named upstream repository. Replace the argument list with the project's real test command before capturing either side of the oracle. Keep markers short and specific so a later reviewer can see them in the saved output without guessing. Do not treat the sample statuses as defaults for every language or every test runner on the project.
{
"command": ["python", "-m", "pytest", "-q", "tests/test_parser.py::test_trailing_comma"],
"before_status": 1,
"after_status": 0,
"failure_marker": "trailing comma",
"success_marker": "1 passed"
}
Capture both sides with one script
The capture script reads the manifest, runs the listed command, and writes status plus both streams into the requested side directory. It does not interpret the bug, and it does not accept a free-form shell string, because argument lists avoid accidental quoting surprises. The contributor runs it once on the untouched checkout, applies the human or model patch, then runs it again with the after side. If the before status already matches the after status, the oracle is invalid and the patch should not be opened.
#!/usr/bin/env python3
"""Proposed local helper. Not a recorded run against an upstream repository."""
import json
import subprocess
import sys
from pathlib import Path
def capture(side: str) -> int:
manifest = json.loads(Path("evidence/manifest.json").read_text())
out = Path("evidence") / side
out.mkdir(parents=True, exist_ok=True)
proc = subprocess.run(
manifest["command"],
text=True,
capture_output=True,
check=False,
)
(out / "stdout.txt").write_text(proc.stdout)
(out / "stderr.txt").write_text(proc.stderr)
(out / "status.txt").write_text(str(proc.returncode))
return proc.returncode
if __name__ == "__main__":
raise SystemExit(capture(sys.argv[1]))
python capture_oracle.py before
python capture_oracle.py after
python check_oracle.py
Check the oracle before review comments exist
The checker loads the manifest and both status files, then confirms the failure marker and the success marker landed on the correct side. It also rejects a success marker that already appears in the before output, which catches tests that never failed. A nonzero checker result means the pull request description must not claim that the captured bug is fixed. The script below is a proposal for a fork-local gate, and it should be adapted to the project's real test entry point.
#!/usr/bin/env python3
"""Proposed checker. Sample markers are not universal pass criteria."""
import json
import sys
from pathlib import Path
def main() -> int:
root = Path("evidence")
manifest = json.loads((root / "manifest.json").read_text())
before_status = int((root / "before" / "status.txt").read_text().strip())
after_status = int((root / "after" / "status.txt").read_text().strip())
before_text = (
(root / "before" / "stdout.txt").read_text()
+ (root / "before" / "stderr.txt").read_text()
)
after_text = (
(root / "after" / "stdout.txt").read_text()
+ (root / "after" / "stderr.txt").read_text()
)
errors = []
if before_status != manifest["before_status"]:
errors.append("before status does not match the manifest")
if after_status != manifest["after_status"]:
errors.append("after status does not match the manifest")
if manifest["failure_marker"] not in before_text:
errors.append("failure marker missing from before output")
if manifest["success_marker"] not in after_text:
errors.append("success marker missing from after output")
if manifest["success_marker"] in before_text:
errors.append("success marker already present before the patch")
if errors:
print("\n".join(errors))
return 1
print("oracle holds: failure captured, then cleared, under one command")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Let a free model review the contract, not invent it
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode's free model access can sit in the review lane after the oracle already exists on disk. The useful task is a bounded reading of the diff against the evidence files, not a fresh verdict on the bug. A careful prompt asks the model to list mismatches, untested branches, and claims that the captured files do not support. The free server option helps when a clean remote run is needed, provided the resulting evidence files still pass the local checker.
Review only evidence/manifest.json, the before and after captures, and the diff.
Do not claim a bug is fixed unless the checker would pass on those files.
List pull request claims that the captured output does not support.
Do not widen the patch, and do not suggest new dependencies.
Split model comments from maintainer decisions
| Question | Model may draft | Human must decide |
|---|---|---|
| Did the before status match the manifest? | Quote the status file | Reject the request if it does not |
| Did the marker move from failure to success? | Point at the two captures | Confirm the marker is specific enough |
| Does the diff touch paths outside the fix and the test? | List the extra paths | Accept the scope or split the series |
| Are license, API, or release notes correct? | Flag uncertainty only | Own the final answer |
| Should this merge? | No draft of a merge vote | Yes, after reading the files |
The table is the review boundary for this workflow, and it stays useful even if the model name changes next month. Current quotas, model lists, and server limits are product facts that expire, so this article does not quote them. Contributors should read the current MonkeyCode documentation before relying on a free run for a time-boxed contribution. A changed limit is a reason to run the same scripts locally, not a reason to skip the oracle.
Keep the patch inside the oracle
A frozen oracle does not prove that every edited line was necessary, so the contributor still reads the diff before opening the request. Extra refactors, renamed helpers, and drive-by formatting changes should move to a separate commit or leave the branch. The model may list those extra paths, but it should not be asked to clean up the module after the evidence is green. A small patch plus a failing-then-passing command is easier to revert than a broad rewrite with a flattering summary.
Commands a reviewer can rerun
The pull request body should contain the manifest command and the checker result, copied from the terminal rather than paraphrased. A reviewer on another machine can then run the same argument list, compare status files, and stop if the marker is too vague. A documented setup target belongs in the manifest notes, and this article does not invent a universal container name. The evidence files stay in the fork until the maintainer asks for them, because some projects prefer not to commit bulky logs.
When the marker lies
A marker that appears in ordinary passing output will make the checker reject a real fix, and that failure is useful. The contributor should choose a string that shows up only in the broken run, such as a specific exception text or a fixture name. If no such string exists, the test itself is the missing artifact, and editing production code first would hide that gap. Flaky commands are out of scope, because a status that changes between two before captures means the oracle is not ready for model review.
Who should skip this gate
This approach is a poor fit for patches that cannot be reduced to one command with a stable marker. Documentation-only edits, visual changes, and rare race conditions need a different proof, because a single captured status can look green while the bug remains. Projects that already require a recorded replay file should follow that project rule instead of adding a second evidence format. The checker also does not replace secret scanning, license review, or a blast-radius look at shared helpers.
Close the branch only after the files agree
A contributor who wants a second reading can paste the checker output and the diff into a free model session. The only requested output should be unsupported claims, checked against files already stored in the evidence directory. Those scripts remain unexecuted proposals, so each fork should adjust markers to its own test runner before trusting a green line. When the two captures and the checker agree, the branch is ready for a human reviewer rather than for another generated summary.
Top comments (0)