DEV Community

Quinn Zhu
Quinn Zhu

Posted on

Gate Your First Agent Patch on a Stable Failure Log

Save the failing command before any agent edits your repo. A first PR should start from a bundle, not a chat guess. That bundle holds a script, two logs, and a base SHA.

Start from the failure

You join the repo and the ticket already sounds urgent. Someone suggests an agent can patch the bug before lunch. You paste the issue text and accept the first diff.

The reviewer cannot replay that chat on their machine. You also cannot defend a rollback if main breaks later. Hour one should prove the failure, not invent the fix.

What you leave behind

You will finish the hour with a small repro directory. It contains a script, two logs, a SHA, and a note. The script must fail the same way on two clean runs.

You do not open a pull request during this hour. You do not let an agent write files before those runs.

Keep the names stable so a reviewer can find them fast. Do not rename logs after you paste one into a prompt.

repro/
  START_SHA
  NOTE
  run.sh
  run-1.log
  run-2.log
  log.diff
  PROMPT_LOG
  run-3.log
Enter fullscreen mode Exit fullscreen mode

1. Freeze the starting SHA

Clone the repository into a new directory on a clean host. Do not reuse a laptop checkout full of old builds. Check out the default branch your team already documents.

Record the commit SHA before you install extra tools. That file is your anchor if the patch needs a revert. Do not update the SHA after you start the two runs.

mkdir -p repro
git rev-parse HEAD | tee repro/START_SHA
printf '%s\n' "ticket: <id>" "command: <paste from CI>" > repro/NOTE
Enter fullscreen mode Exit fullscreen mode

If the branch moves, start the bundle over from clone. A moved base makes the later diff harder to explain.

2. Copy the command, not the story

Open the ticket and the newest failing CI log. Copy the exact command that failed, including its working directory. Do not shorten the command so the run feels faster.

Pin package scripts if CI called those scripts by name. Write unknown flags down instead of guessing their meaning.

cat > repro/run.sh << 'EOF'
#!/usr/bin/env bash
set -euo pipefail
cd "$(dirname "$0")/.."
# Replace the next line with the exact CI command.
npm test -- workspace/parser
EOF
chmod +x repro/run.sh
Enter fullscreen mode Exit fullscreen mode

Treat that file as a template until the real command lands. If the log names a setup step, put that step first. Keep secrets out of the script and out of the note.

3. Run the script twice on a clean host

A laptop cache can hide the bug you are chasing. Run both attempts on a throwaway host instead. A clean host is enough for this narrow check.

Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode offers free model access and a free server option. Use that server only as the clean host for these two runs.

Use free model access only after the logs match. Do not assume a model name, a quota, or a machine size. Those details are not stated here because they can change.

Read the current product docs before you depend on access. Confirm the live limits before you store anything there. Treat current access as optional, not as a lasting promise.

Run the script once and keep every line of output. The true fallback keeps the shell alive after a failed test. You still want the log when the command exits non-zero.

./repro/run.sh > repro/run-1.log 2>&1 || true
Enter fullscreen mode Exit fullscreen mode

Wait a minute, then run the same script again. Save the second log beside the first one. Diff the logs before you decide the failure is real.

./repro/run.sh > repro/run-2.log 2>&1 || true
diff -u repro/run-1.log repro/run-2.log > repro/log.diff || true
Enter fullscreen mode Exit fullscreen mode

Stop when the logs disagree in a meaningful way. A flaky failure is a bad target for your first PR.

Write the mismatch in the ticket and ask a human. Do not ask a model to hide flake with retries.

4. Read the log before you prompt

Read the log from the bottom, not from the top. Find the first error that names a file and a line. Ignore warnings that also appear on a green main build.

If you cannot tie the error to a file, stop here. A vague failure needs a human, not a generated patch.

tail -n 40 repro/run-2.log > repro/PROMPT_LOG
Enter fullscreen mode Exit fullscreen mode

5. Ask for a patch only after the logs match

Open a chat only after the two logs agree. Attach the script, the SHA, and the second log. Ask for a patch that makes this script pass.

Ask for the file list before you ask for code. Reject any reply that ignores the script you saved. Keep the prompt short so a free model can answer.

Trim the saved log before you paste it into a chat. Forty lines is a starting cap, not a measured limit. Never paste tokens, customer rows, or private URLs.

Base SHA: <paste repro/START_SHA>
Command: ./repro/run.sh
Failure log: <paste repro/PROMPT_LOG>
Constraint: change the smallest set of files.
Do not add dependencies.
Return a file list, a patch, and one short reason.
Enter fullscreen mode Exit fullscreen mode

If the ticket needs those, use the private runner. Treat the answer as an unexecuted proposal, not a result. You still have to apply it and rerun the script.

6. Apply it on a branch and rerun

Create a branch from the SHA you already recorded. Apply a short patch by hand so you see each hunk. Do not apply a patch that adds new dependencies.

Run the same script and save a third log. Compare the second log with the third log yourself. The old error text should be gone in the third log.

git checkout -b junior/repro-fix "$(cat repro/START_SHA)"
# Apply the reviewed patch by hand, then:
./repro/run.sh > repro/run-3.log 2>&1 || true
Enter fullscreen mode Exit fullscreen mode

Other tests in that command should not start failing. If they fail, revert the branch and keep the bundle. Map every hunk to a line in the failure log.

git diff --stat "$(cat repro/START_SHA)"
Enter fullscreen mode Exit fullscreen mode

If a hunk has no story, the patch is too wide. Cut the extra files and ask for a smaller patch.

7. Write the PR so a stranger can replay it

Open the PR only when you can narrate the diff. Put the script path and the base SHA in the body.

Name the command you ran and the log you compared. Invite review of the bundle, not praise of the agent.

## Repro
- Base SHA: <paste>
- Script: repro/run.sh
- Compared: repro/run-2.log and repro/run-3.log

## What changed
- <file>: <one sentence tied to the error line>

## Rollback
- Revert this branch if a hunk cannot be explained.
Enter fullscreen mode Exit fullscreen mode

8. Roll back the same day if the story breaks

Roll back the same day if review exposes a gap. A clean revert is better than a clever unexplained patch. Follow your team's revert rule if the merge was squashed.

# Use your team's documented revert path after merge.
git revert --no-edit <merged-commit>
Enter fullscreen mode Exit fullscreen mode

Leave the bundle in the ticket after the revert. The next person can rerun it without your laptop. Do not delete the logs just because the PR closed.

Decision table

Use this table as a gate before you push. It is a checklist, not a scoreboard for speed.

Signal Action
Two logs match and name one error Ask for a small patch
Logs differ, or the run hangs Stop and ask a human about flake
Patch touches files outside the stack Reject it and re-ask with a file cap
Script passes but one hunk has no story Revert and do not merge
Ticket includes secrets or production data Do not use an external host or chat

Print the table and keep it next to the ticket. If a row says stop, stop before you prompt again.

Limits to say out loud

This method fits one failing command with a stable log. It does not fit races, load tests, or hardware bugs. Two matching runs do not prove a rare crash is gone.

A free server may sleep, throttle, or disappear later. Do not call that host your team's continuous integration. Do not store credentials or production dumps on it.

Free model access may reject a very long log. Trim to the failing section and send that slice only. The commands in this article are an unexecuted example.

They are a workflow to adapt, not a measured benchmark. No timings, pass rates, or hardware claims belong here.

Who should skip this

Skip this if your team forbids external code hosts. Skip this if the bug needs a private dataset. Skip this if you are on call for a live incident.

Senior engineers chasing a multi-service outage need another plan. New hires with a documented unit-test failure are the fit. If your mentor already owns the patch, pair with them.

Close the hour

Hour one ends with a bundle, not a hero patch. Save the SHA, the script, and two matching logs.

Ask for a patch only after that evidence exists. If the diff outruns your explanation, revert it today.

If you already have that free server, run the two checks there. Keep the prompt tied to the saved logs only. Check the docs your team trusts before you rely on availability.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to