DEV Community

Roee Hershko
Roee Hershko

Posted on Originally published at levelup.gitconnected.com on

Every Prompt Rule We Added Made Our LLM Agent Worse

Five weeks of measuring a DevOps agent in Slack: what made it shorter, what made it worse, and why the fixes ended up in code

We built Juno, an AI bot that helps engineers in Slack. You ask it things like “why is my deployment stuck?” or “can we add more servers to this service?” It looks at our systems, and when something needs to change, it opens a pull request (a proposed change) for a person to approve.

Every AI bot like this has a system prompt: a page of instructions it reads before every conversation. When Juno gave a bad answer, we did what most teams do. We added a new instruction to that page. The bot got worse. This post is about what we did instead, with the numbers. Written with help from an AI assistant, from my own notes and test results.

One bad answer, one new rule

Juno recommended something we didn’t like, so we added “never headline a recommendation.” It made a change when it should have asked first, so we added “don’t start with today.” It trusted a setting it should have checked, so we added “check the config yourself.” Each new rule fixed the answer it was written for. Then the next bad answer showed up, and we wrote another rule.

After about ten days I was ready to give up. Juno now argued with people, said no to normal requests, and asked people to do work it could do itself. We wanted a helper. We had built a bossy manager.

Why more rules made it worse

Most rules we wrote were “don’t do X.” We meant each one for a single situation. But the model doesn’t apply a rule to one situation. It applies it everywhere. Ten small “don’ts” made Juno careful and defensive about everything.

Then we looked at the logs (the record of what Juno actually did). About half of the bad answers had nothing to do with the instructions. An access key had expired. A tool returned nothing. One part of the system skipped the step that fetches the data. No instruction can fix problems like these.

Three new habits

  1. Keep the instructions short. Say what the job is, not every case. Ours is now under 2,000 characters.
  2. If something must always happen, write code for it. Don’t ask the model nicely. Code doesn’t forget.
  3. When an answer looks wrong, read the logs first. Fix the real problem. If the instructions really are the problem, remove a rule. Don’t add one.

The third habit was hard, even for my AI coding assistant. Every time I showed it one bad answer, it suggested more instructions. At one point I told it: “There might be nothing wrong.” One bad answer doesn’t prove anything.

How we tested

To remove rules safely, we needed to know which ones mattered. So we built tests. Each test is a short conversation, like “my deployment is stuck.” It runs against a fake copy of our systems, so every run sees the same facts. Some checks are simple code: did Juno use the right tool? How many words did it write? Other checks are done by a second AI model, which reads the answer and decides things like “did it ask before making the change?” A full test run cost between a few dollars and a few tens of dollars, which is cheaper than arguing about screenshots.

Two things kept us honest:

  • Run each test several times. AI answers vary. One run proves nothing, so we ran each test six times.
  • Compare on the same day. Models and tools change over time, so we always tested the old version and the new one side by side.

The one sentence that mattered

We removed instructions one group at a time and re-ran the tests after each. Most of them made no difference at all: we deleted them and nothing changed. But one sentence made a big difference:

Read the cluster and the repo yourself with your tools; never ask the person to check something you can read.

Without it, test results dropped from 100 passing out of 102 to 93. In one test, Juno stopped looking at the data and did something else every single time. When we put that one sentence back, results went back to 101 out of 102.

A sentence that only looked guilty

Another sentence told Juno to say “I can’t help” for topics outside its job. We thought it was causing a wrong refusal, so we tested it before deleting it. Without it, Juno made up an answer about a system it knew nothing about in 4 out of 6 runs. With it, 0 out of 6. So the sentence stayed.

The real cause of the refusal was a tool description that didn’t explain what the tool could do. We fixed the description, and the refusal went away. You can’t tell which instruction matters just by reading it. You have to remove it and test.

Asking for short answers made them longer

Juno also wrote too much. Someone asked for a small change. Juno opened the pull request and wrote 168 words, and more than half of them repeated what the pull request already said. Our instructions already said: “Be concise: complete sentences, no filler.” A popular blog post on making Claude less chatty (Opus 5 Made Claude Code Chatty. Three Changes Reined It In.) suggested adding a “give the main point first” section. We tried that, and two variations of it.

That advice worked for the author in a coding tool. For Juno, every version made the repeating worse, and the answers grew from about 71 words to 80, 90 and 108. Our best guess: if you mention the thing you don’t want, the model writes it. It didn’t matter whether we said “do this” or “don’t do that.” Mentioning the pull request text at all made Juno repeat it.

Other things that didn’t work

  • “Reasoning effort” setting. It made answers about a quarter shorter, but never removed a whole paragraph. It controls how hard the model thinks, not how much it writes.
  • Anthropic’s own prompting advice . It helped a little. But when we asked Juno to “briefly explain” something, it still wrote 1,700 to 2,200 characters when the limit was 1,000.
  • A length limit in the instructions. We wrote “max ~1,500 characters.” Every model we tried ignored it.

What worked: deleting text

What helped was removing text. Our documents had about 70 sentences telling Juno how to answer, like “report X” and “start with Y.” The data our tools sent back was full of little explanation notes. The more text the model reads, the more it writes. And instructions about how to answer end up inside the answer.

We deleted those instructions and turned the notes into plain data. Then we re-ran six real questions from the day before. The word counts barely changed. But the answers were clearly better:

  • Gone: “Here’s what I found,” explanations of how to read the data, and headings named after tools.
  • Kept: the real facts, like numbers, times, what was missing and the suggested fix.

So we stopped chasing a lower word count. A short answer that drops the important facts is a worse answer.

One deletion did break something. After we removed “don’t repeat the pull request link,” Juno pasted the link again in every test run. We didn’t put the sentence back. We wrote code that removes links Juno already showed. If a rule was hiding a display problem, replace it with code.

The fix that worked: a size limit in code

Juno still wrote too much. Once it wrote about 170 words when the right answer was “Sure, which service and environment?” So we stopped asking and changed how Juno replies:

  • Juno sends its answer through a “send message” tool.
  • That tool has a hard character limit.
  • The turn ends when the message is sent.

Adding more words to the instructions changed nothing. The limit did all the work. So the only real question is how long you want answers to be. We chose 600 characters, about three lines in Slack.

A surprise side effect

Once “send message” was the last step, Juno changed how it worked. It did everything first, then reported. For a change in production, that meant opening the pull request before the person said yes. In our “ask first, then act” test, the pass rate dropped from 4 out of 6 to 1 out of 6. (The old version already did this about 40% of the time, which was news to us too.)

We also tried another option: let Juno write freely, and if the answer is too long, ask it to shorten it. That got the same length. But once, the shortened answer said Juno had checked something it never checked. Shortening an answer after it’s written can turn “I didn’t check” into “I checked.” We kept the size limit. Next, “ask before changing production” will become code, not an instruction.

The warning we ignored

We added a feature where Juno messages people when their pull request fails a check. Our tests showed these messages were about 200 words long, with code in every one. I thought that was fine and shipped it.

The first real message was correct, but it was a wall of text: about 230 words, two code blocks, and a paragraph about file formats. The useful part, “I can push the fix,” was the very last sentence. When someone didn’t ask for a message, its length is everything. The test showed us the problem, and we ignored it.

What I’d tell anyone building an AI bot

  • If it must always happen, write code. But don’t use code for decisions the model should make. We once wrote code to sort support questions by matching words, and it got the first few real questions wrong.
  • Keep the instructions short. When an answer is wrong, remove a rule. Don’t add one.
  • Don’t mention what you don’t want. Not even as “don’t do this.”
  • Read the logs before changing the instructions. Half our “instruction problems” were broken tools and expired keys.
  • Test on the same day, and run each test several times.
  • Send the model data, not notes. Every explanation you send it, your users will read again.
  • Change the structure, not the wording. A size limit did more than every instruction we wrote. It also changed how Juno worked, so test that too.

We started out writing rules. We ended up deleting most of them, turning a few into code, and keeping one sentence we only knew mattered because we tried removing it.


Top comments (0)