DEV Community

yureki_lab
yureki_lab

Posted on

How I Stopped My AI Coding Agent's Scope Creep With a Diff Budget

TL;DR

My autonomous coding agent kept turning small tasks into sprawling diffs: ask for a bug fix, get a bug fix plus a renamed helper, a reformatted file, and a "quick cleanup" nobody asked for. I fixed it with a diff budget: every task declares which paths it may touch and how big the change is allowed to be, a hook blocks out-of-scope edits before they happen, and a "parking lot" turns the agent's drive-by ideas into follow-up tasks instead of surprise changes. Here's the design, the code, and 5 lessons from running it.

The Problem

I run a fully autonomous implementation system: an orchestrator module hands tasks to parallel implementation agents, and they work through a backlog while I'm doing something else (often sleeping 😴).

The agents are good at the tasks. That was never the problem. The problem was everything they did in addition to the tasks.

A typical example: the task says "fix the off-by-one in the pagination helper." The agent fixes it. Then, since it's already in the file, it:

  • renames two variables "for clarity"
  • converts a callback to async/await
  • notices a similar pattern in a neighboring module and "aligns" it
  • updates a config file because the linter complained about something unrelated

Every one of those changes is defensible on its own. Together they turn a five-line fix into a diff that touches a dozen files, and that hurts in three specific ways:

  1. Review cost explodes. I can verify a five-line fix in a minute. A twelve-file diff where the actual fix is buried somewhere in the middle takes real attention, and attention is the one resource my system can't generate for me.
  2. Parallel agents collide. When several agents work on the same repo, the unrequested edits are exactly the ones that overlap. Two agents never fight over the task they were assigned. They fight over the shared utility file both decided to "tidy up."
  3. Rollbacks get messy. If the fix is wrong, I want to revert the fix. I don't want to also revert an unrelated refactor that some later task now depends on.

My first attempt was the obvious one: add a line to the instructions. "Only change what the task requires. Do not refactor unrelated code."

It helped a little, and not nearly enough. The agent doesn't experience its cleanup as "unrelated." From inside the task, renaming a confusing variable feels like part of doing a good job. I was asking a model to resist something it considers a virtue, using a sentence. That's a losing fight.

Instructions describe intent. If you need a guarantee, you need a mechanism.

So I stopped trying to persuade the agent and built a fence instead.

How I Solved It

The design has three pieces: a budget declared per task, a hook that enforces the path part before an edit lands, and a check on the final diff that enforces the size part. Plus one escape valve, which turned out to be the piece that made everything else work.

flowchart TD
    A[Orchestrator creates task] --> B[Task spec with scope + budget]
    B --> C[Implementation agent starts]
    C --> D{Edit inside allowed paths?}
    D -- yes --> E[Edit applied]
    D -- no --> F[Blocked: agent told to use parking lot]
    F --> G[Parking lot note written]
    E --> H{Final diff within budget?}
    H -- yes --> I[Commit + hand to review]
    H -- no --> J[Returned to agent: trim or request extension]
    G --> K[Orchestrator files follow-up tasks]

1. Every task declares a scope and a budget

When the orchestrator module creates a task, the spec now carries a scope block. It's deliberately boring:

{
  "id": "task-0412",
  "goal": "Fix off-by-one in pagination helper",
  "scope": {
    "allowed_paths": ["src/pagination/**", "tests/pagination/**"],
    "max_files": 4,
    "max_changed_lines": 80
  }
}
Enter fullscreen mode Exit fullscreen mode

Three numbers and a list of globs. The planner step sets them when it writes the task, based on the task type: bug fixes get a tight budget, feature work gets a wider one, and refactors are their own task type with an explicitly large budget. That last part matters. I'm not against refactoring. I'm against refactoring smuggled inside something else.

2. A pre-edit hook blocks out-of-scope writes

Claude Code (I'm on v2.1.x) supports hooks that run before a tool call. A PreToolUse hook receives the tool call as JSON on stdin, and if it exits with code 2, the call is blocked and whatever the hook wrote to stderr is fed back to the model.

That feedback channel is the whole trick. The agent doesn't just hit a wall. It gets told why, and what to do instead.

#!/usr/bin/env python3
"""PreToolUse hook: block edits outside the current task's allowed paths."""
import fnmatch
import json
import os
import sys

call = json.load(sys.stdin)
if call.get("tool_name") not in ("Edit", "Write"):
    sys.exit(0)

with open(os.environ["TASK_SPEC"]) as f:
    scope = json.load(f)["scope"]

root = call.get("cwd", os.getcwd())
target = os.path.relpath(call["tool_input"]["file_path"], root)

if any(fnmatch.fnmatch(target, pattern) for pattern in scope["allowed_paths"]):
    sys.exit(0)

print(
    f"BLOCKED: {target} is outside this task's scope "
    f"({', '.join(scope['allowed_paths'])}).\n"
    "If this change is required to finish the task, stop and request a "
    "scope extension with a one-line reason.\n"
    "If it is an improvement you noticed, append it to the parking lot "
    "file and keep going.",
    file=sys.stderr,
)
sys.exit(2)
Enter fullscreen mode Exit fullscreen mode

Registering it is a few lines in the settings file:

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Edit|Write",
        "hooks": [{ "type": "command", "command": "python3 hooks/scope_guard.py" }]
      }
    ]
  }
}
Enter fullscreen mode Exit fullscreen mode

One honest caveat: this hook only covers the file-editing tools. An agent with shell access can still modify files through a shell command, so the hook is a guardrail, not a security boundary. That's why the next check exists.

3. A post-run check enforces the size budget

Path scoping stops the agent from wandering. It doesn't stop it from rewriting everything inside the allowed directory. So before a task's work is committed, a small script measures the actual diff against the budget:

#!/usr/bin/env python3
"""Compare the working tree diff against the task's budget."""
import fnmatch
import json
import subprocess
import sys

scope = json.load(open(sys.argv[1]))["scope"]

out = subprocess.run(
    ["git", "diff", "--numstat", "HEAD"],
    capture_output=True, text=True, check=True,
).stdout

files, lines, outside = 0, 0, []
for row in out.strip().splitlines():
    added, deleted, path = row.split("\t")
    files += 1
    if added != "-":  # binary files report "-"
        lines += int(added) + int(deleted)
    if not any(fnmatch.fnmatch(path, p) for p in scope["allowed_paths"]):
        outside.append(path)

problems = []
if outside:
    problems.append(f"out-of-scope files: {', '.join(outside)}")
if files > scope["max_files"]:
    problems.append(f"{files} files changed (budget {scope['max_files']})")
if lines > scope["max_changed_lines"]:
    problems.append(f"{lines} lines changed (budget {scope['max_changed_lines']})")

if problems:
    print("OVER BUDGET: " + "; ".join(problems))
    sys.exit(1)
print(f"within budget: {files} files, {lines} lines")
Enter fullscreen mode Exit fullscreen mode

Because this reads the real diff from git, it catches everything regardless of how the file got modified, including the shell-command route the hook can't see.

When the check fails, the task isn't thrown away. The output goes back to the agent with a simple choice: trim the diff down to what the task needs, or request an extension.

4. The parking lot (the piece that made it work)

My first version only had the fence, and it had a nasty side effect: the agent would notice a real problem outside its scope, get blocked, and then the observation simply vanished. Some of those observations were valuable. A genuinely broken neighbor module is something I want to know about.

So each task gets a parking lot: a plain Markdown file where the agent appends things it noticed but wasn't allowed to touch.

## Parking lot: task-0412

- src/search/cursor.py has the same off-by-one pattern as the pagination
  helper. Likely the same bug. Suggested task: bug fix, tight budget.
- Variable names in src/pagination/window.py are misleading (`end` is
  inclusive). Suggested task: refactor, low priority.
Enter fullscreen mode Exit fullscreen mode

When the task finishes, the orchestrator reads the parking lot and files each entry as a candidate task in the backlog, with its own scope and budget. The cleanup still happens. It just happens as its own small, reviewable, revertable change.

This changed the agent's behavior more than the fence did. Once it had somewhere legitimate to put its ideas, it stopped trying to sneak them in.

5. Extensions are allowed, but they cost a sentence

Sometimes the budget is simply wrong. The fix really does need a change to a shared type definition two directories over. For that, the agent can request an extension: it stops, writes one line explaining which path it needs and why, and the orchestrator decides.

Small extensions that stay inside the same package get approved automatically. Anything touching shared or high-traffic areas waits for me. In practice most requests are reasonable, and the ones that aren't are usually a sign the task was specified badly in the first place, which is useful to learn.

Lessons Learned

1. Scope creep is a virtue in the wrong place. The agent isn't misbehaving when it tidies up. It's doing what a conscientious engineer does. Treating it as disobedience leads you to write angrier instructions. Treating it as misrouted effort leads you to build a parking lot.

2. Block with an explanation, never silently. A hook that just says "denied" makes the agent try the same thing a different way. A hook that says "this is out of scope, here are your two options" gets a sensible next move almost every time. The error message is a prompt. Write it like one.

3. Give the agent a legitimate outlet. Every constraint needs an answer to "okay, then what do I do with this?" Without the parking lot, I was throwing away real findings. With it, the constraint became a routing rule instead of a muzzle.

4. Budgets should be set by task type, not by gut. A single global limit is always wrong for somebody: too loose for bug fixes, too tight for features. Three or four task types with different defaults covered nearly everything I needed.

5. Small diffs are a feature of the system, not a nicety. Review speed, parallel agents not colliding, clean reverts: all of these depend on changes being small and single-purpose. If you run agents unattended, diff size is the thing to protect first.

What's Next

  • Smarter budgets. Right now budgets come from static defaults per task type. I want the planner to estimate them from the files the task actually references.
  • Tracking extension requests. If one area of the codebase keeps triggering extensions, that's a hint the module boundaries are wrong there. I'd like that surfaced automatically.
  • Parking lot dedup. Different agents sometimes notice the same issue. Merging those entries before they become tasks is on the list.

Wrap-up

If your coding agent keeps handing you diffs three times bigger than the task, don't write a sterner instruction. Declare a budget, enforce it with a hook, check the real diff, and give the agent a parking lot for its good ideas. You can wire up a first version with the two scripts above in an afternoon.

If this was useful, follow me here on Dev.to. I write regularly about building and running a fully autonomous implementation system with Claude Code. And I'd love to hear in the comments: how do you keep your agent's diffs small? 🚀

Top comments (2)

Collapse
 
reidmarlow profile image
Reid Marlow •

A pre-tool hook on Edit and Write catches direct file edits, while the Bash tool provides an easy bypass. When an agent gets blocked on an edit tool, it often falls back to sed, an echo redirect, or a python one-liner to write the file anyway.

Parsing shell strings with regex gets fragile fast. A cleaner boundary is running git status after each shell command and mounting out-of-scope directories as read-only in the container.

For the line budget, formatters are the main false trigger. An automated Prettier or black pass across an allowed file can turn a five-line fix into a two-hundred-line diff that trips the ceiling. Ignoring whitespace changes in the diff calculation keeps formatting cleanup from exhausting the task allowance.

Collapse
 
xtrel profile image
XtReL | DevSecOps Builder • • Edited

Solid write-up. The line that ties it together for me: "Instructions describe intent. If you need a guarantee, you need a mechanism." I've been writing about the same gap from a metrology angle: a declaration ("the agent says it's done") versus a verified fact. Your post-run check reading the real diff from git is exactly that - it measures what happened, not what the agent reports.

Your "routing rule instead of a muzzle" line is the part I'd underline. The parking lot carries emergent context - things neither you nor the agent knew about before the task started - and gives it a legitimate path instead of the main diff.

One question building on your What's Next: when you start tracking extension requests, do you plan to feed that signal back into the per-type budget defaults, or only use it to flag module boundaries? Those seem like two different fixes - one changes the tolerance, the other changes the system being measured.