An OSS model note stays non-binding until a local contract passes. The contract locks one command, one diff scope, and two exit codes. A free model may comment only after that gate.
Why unbound notes fail
A long model reply can sound like a finished review. It still lacks a rerun of the failing command. Maintainers then merge model confidence instead of local evidence.
This workflow splits the work into three lanes. Local reproduction owns the claim about the bug. A disposable rerun and a model read stay secondary.
Contract fields that matter
Keep the contract in review-contract.json next to the patch. Record only facts that a stranger can check. Leave coding style opinions out of that file.
{
"repo": "example/widget",
"base": "origin/main",
"command": "python -m unittest tests.test_widget",
"expect_before": 1,
"expect_after": 0,
"allowed_paths": ["widget/parse.py", "tests/test_widget.py"],
"model_role": "non_binding_reader"
}
The expect_before field stores the exit before the patch. The expect_after field stores the exit after the patch. Both numbers must come from commands, not from chat.
The widget paths and example.invalid host are placeholders, not a real project. Swap in the public repository before a real review run.
Three lanes, one owner
| Lane | Where it runs | What it proves | Can it approve the patch |
|---|---|---|---|
| Reproduce | Local worktree | Command exit and log | No, it only supplies evidence |
| Rerun | Disposable free server | Same command, clean tree | No, it only confirms the exit |
| Read | Free model session | Risks outside the command | No, comments stay non-binding |
Use the free server as a clean rerun host. Use free model access as a narrow second reader.
The operator presents MonkeyCode as an open-source project. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Those claims cover the rerun lane and the read lane only.
This draft does not name models, quotas, hardware, or duration. Cite a current primary source before publishing any of those figures.
Step 1. Freeze one failing command
Choose a command that fails on the base commit today. Write it into the contract before any edit. Ban prompts, watches, and commands that never exit.
git fetch origin
git switch --detach origin/main
python -m unittest tests.test_widget
echo "before=$?"
Copy the printed exit code into the expect_before field. A zero means this command does not show the bug. Replace that command before writing any code patch.
Run the snippet without errexit, or the shell stops early. A stopped shell never prints the before marker. The missing marker is a failed reproduction, not a pass.
Step 2. Limit the patch to named paths
Open a detached worktree for the patch experiment. Check the diff, then apply it only if the path list matches. A wider diff makes the written contract false.
git worktree add --detach ../review-wt origin/main
git -C ../review-wt apply --check patch.diff
git -C ../review-wt apply patch.diff
git -C ../review-wt diff --name-only > changed.txt
Compare changed.txt with allowed_paths before the next lane. Extra paths mean the contract is already stale. Update the list or remove the extra file.
Step 3. Capture the local after log
Run the locked command inside the review worktree. Save the combined log and the exit marker. Never trim a failure just to save space.
cd ../review-wt
python -m unittest tests.test_widget > after.log 2>&1
echo "after=$?" | tee -a after.log
Require the after marker when expect_after equals zero. A different marker blocks every later review lane. The local log stays the primary review evidence.
Step 4. Confirm on a disposable server
Push nothing private into the disposable server session. Clone the public base, then apply the same patch file. Run the same command and record the server exit beside the local one.
Reuse the same patch.diff on the server lane. Do not rebuild the patch from memory on that host.
# Proposal only. Replace the placeholder before any real run.
git clone --depth 1 https://example.invalid/widget.git review-srv
cd review-srv
git apply --check ../patch.diff
git apply ../patch.diff
python -m unittest tests.test_widget
echo "server=$?"
Label the server block as an unexecuted proposal. A free server may lack compilers, tokens, or long job time. Record that gap and keep the local log as authority.
Step 5. Build a redacted read packet
Copy four items into a folder named packet. Include the contract, the path list, both exit lines, and a scoped diff. Strip secrets, env files, and absolute home paths first.
mkdir -p packet
cp review-contract.json changed.txt packet/
git -C ../review-wt diff -- widget/parse.py tests/test_widget.py > packet/scope.diff
python3 -c 'from pathlib import Path; text = Path("after.log").read_text(encoding="utf-8", errors="replace"); lines = [ln for ln in text.splitlines() if ln.startswith("after=")]; Path("packet/exits.txt").write_text("\n".join(lines) + "\n", encoding="utf-8")'
That packet is the only model input. A full repository upload adds risk without review value. Delete the packet after the note is written.
Step 6. Request a non-binding read
Ask the free model to list uncovered risks. Require each bullet to cite a path or a contract key. Forbid a verdict of correct, safe, or ready.
Role: non_binding_reader
Files: review-contract.json, changed.txt, exits.txt, scope.diff
Task: list risks the locked command does not cover
Rule: cite a path or a contract key in every bullet
Rule: do not approve the patch
A free model session can serve this narrow read. The model can still miss hidden fixtures and local conventions. A human maintainer still decides whether the patch lands.
Step 7. Check the contract before the note is posted
Run the checker on the contract, the path list, and the local log. Post model bullets only when the checker prints success. Prefix each bullet with the non-binding label.
python3 check_review_contract.py review-contract.json changed.txt after.log
The script below is a local unexecuted proposal. It checks fields, role, path scope, and the exit marker. It never contacts a network and it never opens a pull request.
#!/usr/bin/env python3
"""Proposal: validate an OSS review contract before model notes."""
import json
import sys
def main() -> int:
if len(sys.argv) != 4:
print("usage: check_review_contract.py CONTRACT CHANGED LOG")
return 2
contract_path, changed_path, log_path = sys.argv[1:4]
contract = json.loads(open(contract_path, encoding="utf-8").read())
required = [
"command",
"expect_before",
"expect_after",
"allowed_paths",
"model_role",
]
missing = [key for key in required if key not in contract]
if missing:
print("missing fields:", ", ".join(missing))
return 2
if contract["model_role"] != "non_binding_reader":
print("model_role must stay non_binding_reader")
return 2
changed = [
line.strip()
for line in open(changed_path, encoding="utf-8")
if line.strip()
]
extra = sorted(set(changed) - set(contract["allowed_paths"]))
if extra:
print("paths outside contract:", ", ".join(extra))
return 3
log = open(log_path, encoding="utf-8").read()
marker = "after=%s" % contract["expect_after"]
if marker not in log:
print("local log lacks", marker)
return 4
print("contract ok; model notes remain non-binding")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Step 8. File the maintainer note
Use a fixed note shape so reviewers can scan it. Put local evidence first and model text last. Keep the model section short enough to scan.
Command: python -m unittest tests.test_widget
Local before: 1
Local after: 0
Server after: not run
Checker: contract ok
Model role: non-binding
Model notes:
- non-binding: tests/test_widget.py covers only the named parser path
A missing server exit line is still acceptable. A missing local exit line is not acceptable. Readers should rerun the command before they trust the note.
How failed checks should end
Exit 2 means required fields or the model role are wrong. Exit 3 means the diff left the allowed path list. Exit 4 means the local log lacks the expected marker.
Repair the failed lane before another model call. Do not ask a second model to overwrite a bad log. Edit the contract only when the command itself changed.
Limits that stay in force
One command can pass while a neighbor behavior fails. A clean server can still miss private fixtures or system packages. Free model wording can be fluent and still incomplete.
A log line that merely contains the marker can fool this check. Treat the checker as a gate, not as full test proof. Read the unittest summary before accepting the marker.
Normalize relative paths before running the set comparison. A leading prefix can cause a false extra-path failure.
This article states no quota, latency, or uptime figure. Those values were not verified against a primary source here. A free tier is optional assistance, not a promised production host.
Who should not use it
Skip the flow when the patch holds secrets or customer data. Skip it when project rules ban external model reads. Skip it when the bug depends on devices the server cannot attach.
Skip it for broad refactors that no single command represents. Skip it when the review must be a formal security audit. A non-binding model reader cannot replace that audit.
Practice to keep
Store the contract, the log marker, and the checker result with the patch. Mark every model sentence as non-binding in the review thread. Ask one maintainer to rerun the locked command locally.
Optional reruns and model reads can support that second look. They do not replace the local exit code. Operators who keep evidence separate from commentary can try the optional lanes in MonkeyCode.
Top comments (0)