DEV Community

Blake Yang
Blake Yang

Posted on

Freeze a Failing Oracle Before Any Model Touches an OSS Patch

A weekend contributor clones a small library, reproduces a parser bug, and asks a coding model to repair it before dinner. The model returns a tidy diff, a confident summary, and a test that already passes on the patched tree. The maintainer later discovers that the same test also passed on the unpatched tag, so the pull request never proved a fix. That pattern wastes review time, and a frozen oracle stops it before the branch leaves the fork.

Why a passing test is not evidence

An open-source review needs a failure that exists before the patch and disappears after the patch, under the same command. A generated test written after the edit can encode the new behavior without ever having failed. Reviewers then argue about intended behavior instead of comparing two captured runs from the same command. The workflow below treats those two runs as the contract that every later comment must respect.

Run the loop in order

  1. Write the manifest with an argument list, both expected statuses, and two markers before any production edit.
  2. Capture the before side on the untouched checkout and stop if the status is already the success code.
  3. Apply the smallest patch that can clear the failure marker, whether a person or a model drafted the diff.
  4. Capture the after side, run the checker, and refuse the pull request text when the checker returns an error.
  5. Ask the model to compare the diff with those files, and leave the merge decision with a human reviewer.

Lay out the evidence before editing

The contributor creates an evidence directory before touching production code, then records the failing command in a small manifest. That manifest names the command, the expected nonzero status, and a short marker that must appear in the failure output. A later checker refuses to bless a patch when either capture is missing or when the command string has changed. The point is narrow: prove this bug, on this command, before any model rewrites the story.

Directory contract

evidence/
  manifest.json
  before/stdout.txt
  before/stderr.txt
  before/status.txt
  after/stdout.txt
  after/stderr.txt
  after/status.txt
Enter fullscreen mode Exit fullscreen mode

Proposed manifest

The sample below is a proposal for one pytest node, and it is not a result from a named upstream repository. Replace the argument list with the project's real test command before capturing either side of the oracle. Keep markers short and specific so a later reviewer can see them in the saved output without guessing. Do not treat the sample statuses as defaults for every language or every test runner on the project.

{
  "command": ["python", "-m", "pytest", "-q", "tests/test_parser.py::test_trailing_comma"],
  "before_status": 1,
  "after_status": 0,
  "failure_marker": "trailing comma",
  "success_marker": "1 passed"
}
Enter fullscreen mode Exit fullscreen mode

Capture both sides with one script

The capture script reads the manifest, runs the listed command, and writes status plus both streams into the requested side directory. It does not interpret the bug, and it does not accept a free-form shell string, because argument lists avoid accidental quoting surprises. The contributor runs it once on the untouched checkout, applies the human or model patch, then runs it again with the after side. If the before status already matches the after status, the oracle is invalid and the patch should not be opened.

#!/usr/bin/env python3
"""Proposed local helper. Not a recorded run against an upstream repository."""
import json
import subprocess
import sys
from pathlib import Path

def capture(side: str) -> int:
    manifest = json.loads(Path("evidence/manifest.json").read_text())
    out = Path("evidence") / side
    out.mkdir(parents=True, exist_ok=True)
    proc = subprocess.run(
        manifest["command"],
        text=True,
        capture_output=True,
        check=False,
    )
    (out / "stdout.txt").write_text(proc.stdout)
    (out / "stderr.txt").write_text(proc.stderr)
    (out / "status.txt").write_text(str(proc.returncode))
    return proc.returncode

if __name__ == "__main__":
    raise SystemExit(capture(sys.argv[1]))
Enter fullscreen mode Exit fullscreen mode
python capture_oracle.py before
python capture_oracle.py after
python check_oracle.py
Enter fullscreen mode Exit fullscreen mode

Check the oracle before review comments exist

The checker loads the manifest and both status files, then confirms the failure marker and the success marker landed on the correct side. It also rejects a success marker that already appears in the before output, which catches tests that never failed. A nonzero checker result means the pull request description must not claim that the captured bug is fixed. The script below is a proposal for a fork-local gate, and it should be adapted to the project's real test entry point.

#!/usr/bin/env python3
"""Proposed checker. Sample markers are not universal pass criteria."""
import json
import sys
from pathlib import Path

def main() -> int:
    root = Path("evidence")
    manifest = json.loads((root / "manifest.json").read_text())
    before_status = int((root / "before" / "status.txt").read_text().strip())
    after_status = int((root / "after" / "status.txt").read_text().strip())
    before_text = (
        (root / "before" / "stdout.txt").read_text()
        + (root / "before" / "stderr.txt").read_text()
    )
    after_text = (
        (root / "after" / "stdout.txt").read_text()
        + (root / "after" / "stderr.txt").read_text()
    )
    errors = []
    if before_status != manifest["before_status"]:
        errors.append("before status does not match the manifest")
    if after_status != manifest["after_status"]:
        errors.append("after status does not match the manifest")
    if manifest["failure_marker"] not in before_text:
        errors.append("failure marker missing from before output")
    if manifest["success_marker"] not in after_text:
        errors.append("success marker missing from after output")
    if manifest["success_marker"] in before_text:
        errors.append("success marker already present before the patch")
    if errors:
        print("\n".join(errors))
        return 1
    print("oracle holds: failure captured, then cleared, under one command")
    return 0

if __name__ == "__main__":
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

Let a free model review the contract, not invent it

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode's free model access can sit in the review lane after the oracle already exists on disk. The useful task is a bounded reading of the diff against the evidence files, not a fresh verdict on the bug. A careful prompt asks the model to list mismatches, untested branches, and claims that the captured files do not support. The free server option helps when a clean remote run is needed, provided the resulting evidence files still pass the local checker.

Review only evidence/manifest.json, the before and after captures, and the diff.
Do not claim a bug is fixed unless the checker would pass on those files.
List pull request claims that the captured output does not support.
Do not widen the patch, and do not suggest new dependencies.
Enter fullscreen mode Exit fullscreen mode

Split model comments from maintainer decisions

Question Model may draft Human must decide
Did the before status match the manifest? Quote the status file Reject the request if it does not
Did the marker move from failure to success? Point at the two captures Confirm the marker is specific enough
Does the diff touch paths outside the fix and the test? List the extra paths Accept the scope or split the series
Are license, API, or release notes correct? Flag uncertainty only Own the final answer
Should this merge? No draft of a merge vote Yes, after reading the files

The table is the review boundary for this workflow, and it stays useful even if the model name changes next month. Current quotas, model lists, and server limits are product facts that expire, so this article does not quote them. Contributors should read the current MonkeyCode documentation before relying on a free run for a time-boxed contribution. A changed limit is a reason to run the same scripts locally, not a reason to skip the oracle.

Keep the patch inside the oracle

A frozen oracle does not prove that every edited line was necessary, so the contributor still reads the diff before opening the request. Extra refactors, renamed helpers, and drive-by formatting changes should move to a separate commit or leave the branch. The model may list those extra paths, but it should not be asked to clean up the module after the evidence is green. A small patch plus a failing-then-passing command is easier to revert than a broad rewrite with a flattering summary.

Commands a reviewer can rerun

The pull request body should contain the manifest command and the checker result, copied from the terminal rather than paraphrased. A reviewer on another machine can then run the same argument list, compare status files, and stop if the marker is too vague. A documented setup target belongs in the manifest notes, and this article does not invent a universal container name. The evidence files stay in the fork until the maintainer asks for them, because some projects prefer not to commit bulky logs.

When the marker lies

A marker that appears in ordinary passing output will make the checker reject a real fix, and that failure is useful. The contributor should choose a string that shows up only in the broken run, such as a specific exception text or a fixture name. If no such string exists, the test itself is the missing artifact, and editing production code first would hide that gap. Flaky commands are out of scope, because a status that changes between two before captures means the oracle is not ready for model review.

Who should skip this gate

This approach is a poor fit for patches that cannot be reduced to one command with a stable marker. Documentation-only edits, visual changes, and rare race conditions need a different proof, because a single captured status can look green while the bug remains. Projects that already require a recorded replay file should follow that project rule instead of adding a second evidence format. The checker also does not replace secret scanning, license review, or a blast-radius look at shared helpers.

Close the branch only after the files agree

A contributor who wants a second reading can paste the checker output and the diff into a free model session. The only requested output should be unsupported claims, checked against files already stored in the evidence directory. Those scripts remain unexecuted proposals, so each fork should adjust markers to its own test runner before trusting a green line. When the two captures and the checker agree, the branch is ready for a human reviewer rather than for another generated summary.

Top comments (0)