DEV Community

Cover image for The agent did exactly what I asked. That was the bug.
alphanumericentity
alphanumericentity

Posted on

The agent did exactly what I asked. That was the bug.

Three times this month my agent executed a flawed request perfectly. Hooks can't catch this one, because the mistake isn't in any single tool call, it's in the plan. Here's what I changed instead.

I wrote a while back about moving rules out of my instructions file and into hooks, because a rule in prose gets followed most of the time and a rule in a tripwire gets followed every time. I still believe that. Someone in the comments pushed on where determinism stops being the right tool, and I didn't have a good answer then.

I do now, because three times in the past month my agent did precisely what I asked, did it well, and the result was wrong anyway. Not a hallucination, not a bad diff. Correct execution of a flawed request. That class of failure has no tool call to intercept, which makes it the exact thing a hook cannot save you from.

One: I asked for a demo GIF and got a security bug

I have a side project that redacts PDFs in the browser. I wanted an animated demo for the landing page, so I asked the agent to drive the app with Playwright and record the whole flow: open a file, drag a box over a name, export, done.

It built the script. The script worked. I watched the GIF and the black box was sitting next to the name instead of on top of it.

The redaction code was placing boxes by splitting a text run's total width by its character count, so a box over characters 4 to 9 was a flat fraction of the run. That's correct for a monospaced font and wrong for every other font, and the drift grows with the length of the run. Sometimes it covered the word. Sometimes it covered the word and a bit of the next one. Sometimes it sat just to the right of the thing I was trying to hide.

Here is the part that stings. The app already verified its own output. After exporting, it re-opens the PDF it just produced, extracts the text layer and checks that the redacted string is gone. That check passed every single time, and it was right to pass: the text really was destroyed. It just wasn't the text under the box the user drew. My verifier and my renderer were computing position the same wrong way, so the verifier could not see the error even in principle.

I asked for a marketing asset. What found the bug was that a GIF is looked at by a human with eyes, and every assertion I had written shared an assumption with the code it was testing.

Two: a perfect entry in the wrong section

Different day. I wanted my project listed in one of the big awesome-privacy lists, so I asked the agent to add it to the Office section and open the PR.

It did the work properly. Read the contributing guide, matched the entry format down to the trailing source-link convention the other entries use, filled in the seven-point checklist, picked an insertion point that wouldn't conflict with other open PRs, wrote a disclosure paragraph. Genuinely good work.

I caught it at the last second, with the commit dialog open: that section is titled "Avoid Microsoft Office and Google Docs, instead use", and everything in it is a suite, an editor or a spreadsheet. A PDF utility in there is a category error. Any maintainer would have closed it, and fairly.

The agent never had a reason to question the section. I named it. It optimized for "is this entry well-formed and well-argued", which it was, and never for "is Office the right section", because I'd already answered that.

Three: a confident click path to a setting that lives somewhere else

Smallest one, same shape. I asked where a particular GitHub repo setting lives. I got "Settings, then General", stated plainly, with the surrounding steps.

It isn't there. That field is behind a gear icon next to the About panel on the repo's front page, which is not Settings at all. The answer was assembled from what a settings page plausibly contains. When I made it go read the documentation page first, it came back with the right path in one shot, and with the character limits and naming rules I hadn't thought to ask about.

What these have in common

The agent is a premise-taker. It will interrogate its own implementation all day and will not touch the frame of the request, because the frame came from you and you're the one it's trying to satisfy. "Did I do what was asked" is a question it can answer. "Was the ask right" isn't even in scope.

Four things I changed, in rough order of how much they've paid off.

Ask for the decision, not the destination. "Find where this belongs in the list, then add it" instead of "add it to the Office section". One extra clause moves the category judgement inside the task, where the agent will actually spend reasoning on it. When I name the destination, I've silently converted a decision into a constraint.

Primary source, or say it's unverified. For anything about a UI, a config field, an API shape, a provider's semantics: read it this session and cite it, or label the answer as a guess. Plausible-sounding click paths are the most confident wrong answers I get, because nothing in the model's experience punishes a well-formed guess about a settings page.

Make it tell me what it assumed. Every non-trivial task ends with what it had to assume and what it could not check. This is cheap and it has caught more than I expected, because the premise shows up in writing where I can see it, instead of staying implicit in a diff that looks fine.

Build the artifact that can disagree with me. A screenshot, a recording, a pixel diff, a second measurement taken a different way. The GIF found the redaction bug for one reason: it didn't share any code with the thing it was checking. Now that project has a test that pins box positions against independently computed glyph metrics and then samples the burned pixels, which is the same idea with the human taken out of the loop. It's here if you want the shape of it, AGPL like the rest of it. If your check is built from the same assumption as the code it checks, you don't have a check, you have a second copy of your belief.

Why a hook can't help here

I tried to find the tripwire for this and there isn't one, which took me a while to accept.

A hook fires on a single tool call and decides from that call's input alone. That works when the mistake is local: a dangerous shell command, a comment block, an edit to a file that should never change. A wrong premise doesn't look like that. It produces a sequence of individually reasonable calls, each one defensible on its own, all of them marching toward a thing that shouldn't have been built. There's no string to match, no file path to guard, no moment where the input looks wrong.

Determinism needs a decidable question. "Is the Office section the right home for a PDF tool" is not one.

So the fix isn't mechanical. It's that the review has to happen earlier than I was doing it, before the work rather than on the diff. For agent work I'd now go further: reviewing the plan is worth more than reviewing the code, because the code is usually right. The expensive errors arrive fully formed, well-tested and pointed in the wrong direction.

That's the part I'd like to be argued with on. Do you actually review plans before the agent starts, and does it hold up, or does it just move the rubber stamp one step earlier? Because I notice I'm much better at approving a plan I half-read than at approving a diff I half-read, and I don't think that's a good sign.

Top comments (0)