DEV Community

prodbymarcu
prodbymarcu

Posted on

My local LLM fixed a bug overnight. The "success" was the bug.

I run a small autonomous coding agent on a secondhand RTX 3090 in my apartment. Last week I pointed it at paid GitHub bounties and let it run unsupervised, thinking I'd wake up to pull requests. I did wake up to pull requests. One of them was a lie, and it taught me more about agent pipelines than a week of everything going right.

The model is Qwen3 27B, quantized, served locally with llama.cpp. It is not a frontier model. That's the whole point. Frontier models can write correct diffs most of the time. A 27B model will happily produce something that looks like a diff, fails to apply, and tells no one. Building for that gap is the actual engineering.

The setup

The pipeline is boring on purpose. A Python runner loops over bounties, clones the repo, installs dependencies, runs the test suite to get a baseline, then asks the model to fix the issue in up to six rounds. Each round: generate a patch, apply it, run tests, feed failures back. If tests pass, save the patch for human review.

The naive version of "apply the patch" is:

applied = subprocess.run(["git", "apply", "-"], input=diff, ...)
run_tests()
Enter fullscreen mode Exit fullscreen mode

That has a flaw that took me an embarrassing amount of logging to see.

The false green

Round one on a real bounty came back green in 40 seconds. One round, tests pass, done. I checked the diff file and it was empty. Zero bytes.

What happened: the model produced a diff with fake hunk headers. git apply rejected it with "corrupt patch", which my code caught. It logged the failure and moved to the next round. But the test suite runs on the repo either way, and the repo was in its original state. Tests passed because the original code was fine. My pipeline saw exit code 0 and declared victory over a patch that never existed.

I'd built a machine that manufactures green checkmarks out of nothing. If I hadn't looked at the diff, that empty file would have gone straight to a maintainer under my name.

The fix is two lines of paranoia:

changed = subprocess.run(
    ["git", "status", "--porcelain"], cwd=repo, capture_output=True, text=True
).stdout.strip()
if not changed:
    continue  # patch never applied; tests passing means nothing
Enter fullscreen mode Exit fullscreen mode

If the working tree is clean after "applying" a patch, the patch did not apply. Any test result you get after that is a test of the pristine code, and your pipeline will happily attribute the pristine code's success to a patch that isn't there.

The subtle cousin of that bug

There's a sneakier version. git apply --3way stages successful merges into the index. Plain git diff shows only unstaged changes, so a successful three-way apply can leave you staring at an empty diff even though the patch landed. Use git diff HEAD when you're harvesting the final patch for review, not git diff. Same family of bug: your verification step looking at the wrong layer of git state.

Bloat is also a failure mode

Once the apply path was honest, a different problem surfaced. The model, when asked to fix a one-line pagination bug in a 58KB React context file, rewrote the whole file. The rewrite typechecked. It also dropped a localStorage token refresh and swapped a PATCH endpoint for a POST against a different URL path. Technically valid TypeScript, functionally a regression.

The diff was 1,379 lines. The bug was one line.

Small models don't write surgical patches when given a whole file. They echo the file back with their edit smeared through it, and the echo drifts. So the runner now rejects any patch over 60 changed lines for a bounty that describes a focused bug. If the model can't express the fix in under 60 lines, it doesn't understand the fix, and I don't want it near my GitHub account.

What actually worked

Two things fixed the rewrite problem.

First, stop asking for diffs. My model produces valid unified diffs maybe never. Instead I give it the exact file contents and ask for the complete new file back, or for big files, a SEARCH/REPLACE block with lines it must copy character for character. Then I do the replacement in Python and verify the SEARCH text actually exists in the file. If it doesn't, the round fails loudly instead of silently.

Second, send the whole file or don't send it. I initially capped file context at 20K characters to save tokens. The target file was 58KB. The model invented the back half of the file from vibes, and the invention compiled. Truncation doesn't make a small model careful. It makes it confident and wrong with valid syntax.

The part where I actually got a PR out of it

After all that, the same pipeline produced two real patches, both verified against the repos' own typecheck with no new errors: a pagination state bug where hasMoreBounties was computed from a stale closure value, and an unauthenticated DoS in a QR code endpoint where the size query param went straight into a PNG buffer allocation with no upper bound. Small fixes, real issues, patches under 25 lines each.

The pipeline that produced those is the same one that produced the empty diff and the 1,379-line rewrite. The difference is a handful of guards that assume the model is wrong until the filesystem proves otherwise.

If you're wiring a local model into anything autonomous, spend your paranoia budget on the verification path, not the prompt. The prompt is the easy part. The part that checks whether anything actually happened is where your nights get saved or wasted.

One last thing for anyone running this against real repos under their own name: review every diff yourself before it ships. The guard rails catch the bugs you anticipated. The maintainer reading your PR catches the ones you didn't. Better to have that conversation over a good patch than a fabricated one.

Top comments (0)