A model review is useful only after a frozen replay packet exists. That packet stores the pinned commit, the failing command, and the patch. The model may write dissent notes, but the local oracle decides the merge.
Open-source fixes fail when review context lives only in a chat window. A later reader cannot replay the bug, the diff, or the exact prompt. A review packet turns that context into files with a recorded result.
Start from a pinned failure
Contribution quality drops when the patch and the bug drift apart. The author pins the upstream commit before any edit or model call. The manifest records the clone URL, the commit, and the test command.
This step is not a fixture shrink and not a path allowlist. Those earlier bounds limit the files under edit. This packet bounds the evidence a later reviewer can replay.
What belongs in the packet
The packet stays small enough to copy onto a clean server. It stores inputs, commands, and the expected local result. Chat transcripts and screenshots stay out of the tree.
review-packet/
manifest.json
repro.sh
fix.patch
fail.out
fail.err
fail.exit
notes/dissent.md
notes/scored.json
oracle/checklist.txt
score.py
The manifest names the project, the commit, and the test command. It does not name a model, a quota, or a hardware size. Those fields change, and a packet should not freeze unverified claims.
Run seven packet steps
The author follows these steps in order on a clean worktree. The workflow stops when the first required artifact is missing. A skipped step makes the later model note unverifiable.
1. Pin the upstream commit
The author clones the repository and checks out a known commit. The workflow refuses to continue when the worktree is dirty. A dirty tree hides unrelated edits inside the review.
git clone "$UPSTREAM_URL" repo
cd repo
git fetch origin "$PINNED_SHA"
git checkout --detach "$PINNED_SHA"
test -z "$(git status --porcelain)"
The status test must succeed before the packet proceeds. A shallow clone can save time on a large history. Full history is fetched when blame or bisect becomes necessary.
2. Capture the failing command
The author writes one command that fails for the reported bug. The capture saves stdout, stderr, and the exit code beside the manifest. The author does not patch the tree during this capture step.
The author places the failing command in repro.sh before any edit. The script below is a proposal and was not executed here. The script points at one reported failure, not the whole suite.
#!/usr/bin/env bash
set -euo pipefail
pytest tests/test_parser.py::test_empty -q
The author saves streams after that script runs on the pinned commit. The exit code stays in a separate file for later comparison. The capture sequence below is proposed and was not executed here.
set -o pipefail
bash repro.sh > fail.out 2> fail.err
echo $? > fail.exit
test "$(cat fail.exit)" -ne 0
A zero exit means the bug is not reproduced yet. The author stops and fixes the command before any model call. A model cannot validate a failure the packet never shows.
3. Keep the human patch minimal
The author applies the smallest change that targets the captured failure. The author commits it locally with a message that names the bug. Generated suggestions stay in a side file, not in the diff.
git add -p
git commit -m "Fix report parser on empty input"
git format-patch -1 --stdout > fix.patch
git apply --check fix.patch
The patch file is the only code a reviewer should read first. Extra refactors belong in a later commit with a separate packet. This keeps blame, revert, and review inside one bounded change.
4. Write the manifest
The manifest records what the author claims and what the reviewer can rerun. The author fills every field before a model prompt is sent. The JSON below is a proposal with placeholders, not a measured run.
{
"project": "example-cli",
"upstream": "https://example.invalid/org/example-cli.git",
"commit": "PINNED_SHA",
"fail_command": "pytest tests/test_parser.py::test_empty -q",
"pass_command": "pytest tests/test_parser.py -q",
"patch": "fix.patch",
"oracle": "oracle/checklist.txt"
}
A real upstream URL belongs only after an operator verifies it. The example host above is intentionally not a live project. The author replaces every placeholder before the packet leaves the machine.
5. Replay on a clean server
A clean server matters when the laptop cannot host the toolchain. MonkeyCode, per the operator, offers free model access for the later dissent pass. The same source reports a free server option for this replay.
Disclosure: This article was prepared as part of MonkeyCode's product outreach. Current quotas, model names, and duration are not stated here. The operator confirms those terms on the product page before a long run.
The packet should still run on any clean host with the same commands. Those commands, not the vendor, define the replay. The manifest stays free of quota claims and hardware guesses.
The author copies the packet, not the chat history, onto that server. The server runs the fail command on the pinned commit before the patch. The patch is applied only after the saved failure code matches.
python3 score.py check-fail manifest.json
git apply --check fix.patch
git apply fix.patch
python3 score.py check-pass manifest.json
A pass after the patch is necessary, but it is not sufficient. Hidden regressions still need the broader command named in the manifest. Both exit codes are recorded in the packet before any model call.
6. Ask only for dissent
The author sends the manifest, the patch, and the oracle checklist. The prompt asks for disagreements, missing tests, and risky assumptions. The prompt does not request a rewrite or a merge approval.
A bounded prompt keeps the review inside the packet. The prompt below is proposed text and was not executed for this draft. The author saves the reply as notes/dissent.md with a UTC timestamp.
The reviewer has no merge rights.
Read manifest.json, fix.patch, and oracle/checklist.txt.
List dissent notes only.
For each note, cite a file and a reason.
Mark confidence as low, medium, or high.
Do not propose a replacement patch.
Do not claim a test ran unless the packet shows an exit code.
The author never pastes the reply over the patch without a human diff. The dissent file is evidence, not an instruction to edit. A confident tone does not raise the note above the oracle.
7. Score notes against the oracle
The oracle is a local checklist the author already trusts. The author writes that checklist from the bug report and failing command. A checklist written after the model reply is not an oracle.
empty input returns a typed error, not a stack trace
parser does not read paths outside the fixture directory
fail command exits non-zero on the pinned commit
patch does not touch lockfiles or generated assets
Each dissent note becomes confirmed, rejected, or untested. Untested notes block the pull request until a command exists. The sample score file below is a proposal, not a recorded review.
{
"notes": [
{
"id": "D1",
"file": "src/parser.py",
"reason": "empty input still falls through to the generic handler",
"confidence": "medium",
"status": "untested"
}
]
}
The scorer below is a proposal, not a benchmark. It checks status labels and does not rank models. It also does not measure speed or product uptime.
STATUSES = {"confirmed", "rejected", "untested"}
def gate(notes):
unknown = [n for n in notes if n.get("status") not in STATUSES]
if unknown:
raise SystemExit("unknown status")
pending = [n for n in notes if n["status"] == "untested"]
if pending:
raise SystemExit("untested dissent remains")
return "ready-for-human-review"
A green gate means the packet is reviewable, not that the patch is correct. A wrong human label can still pass this script. The method catches a missing review, not a mistaken judgment.
Apply one decision row
The author chooses one row before opening a pull request. The table is a rule set, not a measured study. Empty dissent never counts as model or maintainer approval.
| Local result | Model note | Next action |
|---|---|---|
| Failure not reproduced | Any note | Repair the packet first |
| Failure reproduced, patch still fails | Any note | Repair the patch locally |
| Patch passes and note is confirmed | Cites a real gap | Add a test or narrow the claim |
| Patch passes and note is rejected | Contradicts the oracle | Record the rejection reason |
| Patch passes and note is untested | No command yet | Hold the pull request |
| Patch passes and dissent is empty | No notes | Human review remains required |
A silent model can miss a race, a permission check, or a bad default. Human review stays mandatory after the table selects a row. The pull request should link the row that was chosen.
Limitations to state in the request
The packet cannot see private maintainer context or unpublished issues. A free server may lack services, secrets, or hardware the bug needs. The author does not copy credentials into the packet or the model prompt.
Free model access can change in quota, model choice, or availability. This article states no token count, no uptime, and no speed claim. The operator verifies the live terms before relying on a free run.
Generated dissent can look precise while citing a stale line. The author compares every citation to the patch before trusting it. The author discards notes that name files absent from the packet.
Who should skip this method
This method is a poor fit for a one-line typo with an obvious test. The packet overhead is larger than that change. A direct pull request is enough in that narrow case.
This method is a poor fit when the bug needs secrets or private logs. A free server and a model prompt are the wrong place for those inputs. The author redacts first or keeps the work on a trusted machine.
This method is a poor fit when no local oracle exists for the bug. A model note without a checklist becomes an unreviewed opinion. The author builds the checklist before spending a model call.
Close the loop in the pull request
The pull request links the packet, the dissent file, and the score output. The body states which notes were confirmed and which were rejected. Maintainers can replay the commands without reading a chat log.
The body names the pinned commit in the pull request text. The body attaches the fail exit code and the pass exit code. That record outlives the chat session that produced the dissent.
When a clean host is missing, the operator checks current MonkeyCode terms. The author runs this same packet once under the free server option. That run counts as a replay, not as universal proof.
The local oracle and the human reviewer still own the merge. A free run does not replace the pinned failure or the dissent score. The author publishes the packet with the pull request for later replay.
Top comments (0)