<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bryan Gelfius</title>
    <description>The latest articles on DEV Community by Bryan Gelfius (@bgelfius).</description>
    <link>https://dev.to/bgelfius</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F718867%2F1a1609bd-f06a-4b1e-9b02-39927ca1d62b.jpeg</url>
      <title>DEV Community: Bryan Gelfius</title>
      <link>https://dev.to/bgelfius</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bgelfius"/>
    <language>en</language>
    <item>
      <title>Five Ways My AI Agents Went Wrong (and What I Changed)</title>
      <dc:creator>Bryan Gelfius</dc:creator>
      <pubDate>Wed, 07 Oct 2026 18:19:05 +0000</pubDate>
      <link>https://dev.to/bgelfius/five-ways-my-ai-agents-went-wrong-and-what-i-changed-6gk</link>
      <guid>https://dev.to/bgelfius/five-ways-my-ai-agents-went-wrong-and-what-i-changed-6gk</guid>
      <description>&lt;p&gt;I run a small software company by myself. Claude Code agents play the engineers, QA, marketing and so on. It works better than I expected, but almost every problem I've had came from the same place: I believed something I hadn't checked. An agent's "done," a file the agents read, a doc that looked current.&lt;/p&gt;

&lt;p&gt;Here are five times that bit me, and what I do now. At the end there's a small free toolkit with the guardrails, if you'd rather not rebuild them.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. A persona file said something false, and every agent repeated it
&lt;/h2&gt;

&lt;p&gt;My agents each have a role file. One of them said, flat out, that a certain ad platform bans a certain kind of ad. I never tested that. Nobody did. I traced it back through &lt;code&gt;git log -p --follow&lt;/code&gt; and it showed up in the file's first commit as a guess.&lt;/p&gt;

&lt;p&gt;After that it got cited in review after review. Each review quoted the one before it, not the platform's policy. It got as far as influencing whether I'd rename something. When I finally read the actual policy, the real rule was much narrower than what my file claimed.&lt;/p&gt;

&lt;p&gt;Now anything in a role file about the outside world (platform rules, vendor limits, marketplace requirements) counts as a guess until someone points me to the source. And when I find one that's wrong I fix the role file itself, not just the one review, or it comes back next session.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. A background agent made up a confirmation
&lt;/h2&gt;

&lt;p&gt;I asked an agent to draft listing copy for one product and told it to leave everything else alone. It said it was done and only mentioned the one file.&lt;/p&gt;

&lt;p&gt;I ran &lt;code&gt;git diff&lt;/code&gt; anyway. It had edited four other files, and in one of them it had written that I'd confirmed a listing was live. I never said that. It was invented, and it was sitting in a checklist that later sessions would have taken as fact.&lt;/p&gt;

&lt;p&gt;So now, every time an agent says it's finished:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git status
git diff &lt;span class="nt"&gt;--stat&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the whole repo, not just the files I assigned. Out-of-scope edits get reverted, not cherry-picked. And if agent output says "the user confirmed X," I don't believe it unless X is actually in the conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. A script I forgot about wiped out my hand-written notes
&lt;/h2&gt;

&lt;p&gt;An agent was refreshing some status data and ran one of my own scripts from an &lt;code&gt;integrations/&lt;/code&gt; folder. The script regenerated the whole output file from scratch. That file also had months of my own analysis in it. One run replaced all of it with a bare machine-generated stub, a 202-line change, and it brought back some stale data I'd specifically marked as unreliable.&lt;/p&gt;

&lt;p&gt;I only caught it because &lt;code&gt;git diff --stat&lt;/code&gt; showed a big change in a file the task had nothing to do with. I reverted before committing.&lt;/p&gt;

&lt;p&gt;The script is append-only now. It adds a new dated block under a known heading and never deletes or replaces a line. If it can't find the heading, it stops instead of falling back to overwriting. I tested it with a dry run and compared it line by line against the real file. The bigger lesson: before you run any script, or let an agent run it, find out which files it writes and whether any of them hold something a script can't recreate.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Two sessions, one working tree
&lt;/h2&gt;

&lt;p&gt;I usually have a few Claude Code sessions open on the same clone. They share one index and one working directory. One session ran &lt;code&gt;git add -A&lt;/code&gt; and committed, and it swept up a multi-file edit another session had made but not committed yet. The content was right. It just landed under an unrelated commit message, mixed in with other work.&lt;/p&gt;

&lt;p&gt;Another time a session stashed someone else's uncommitted edits and switched branches. Nothing was lost, but the files I was editing vanished out from under me.&lt;/p&gt;

&lt;p&gt;What I do now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Commit your own edits in the same turn you make them.&lt;/li&gt;
&lt;li&gt;Stage explicit paths, &lt;code&gt;git add path/to/file&lt;/code&gt;, never &lt;code&gt;git add -A&lt;/code&gt; in a shared clone.&lt;/li&gt;
&lt;li&gt;For anything longer than a quick turn, use a worktree: &lt;code&gt;git worktree add ../repo-topic some-branch&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. My status docs were wrong in both directions
&lt;/h2&gt;

&lt;p&gt;I keep a priorities doc and a design doc. During one push, they said a feature had shipped while &lt;code&gt;gh pr view&lt;/code&gt; showed the PR still open. A few hours later it really had merged, so things had just moved. In the same session I thought three tickets were done because the priorities doc said so. I read the code and found no trace of any of them.&lt;/p&gt;

&lt;p&gt;So the doc was ahead of reality in one place and behind in another. Neither the doc nor my memory of it counted as evidence.&lt;/p&gt;

&lt;p&gt;Now "done" means three things I check myself: a merged PR (&lt;code&gt;gh pr view&lt;/code&gt;), merged code that isn't a stub, and a deploy that actually happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  What these have in common
&lt;/h2&gt;

&lt;p&gt;None of this was the model being dumb. Every time, something got treated as verified when it was only plausible: a persona file, an agent's own summary, a script I'd forgotten, a status doc. The fix is always the same. Check the thing that actually knows the truth (the diff, the PR, the primary source, the real file) before acting.&lt;/p&gt;

&lt;p&gt;I apply the same idea to releases. Nothing ships without two separate approvals, and neither one can approve for the other. The last gate in my release playbook looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;QA APPROVED       [date] [who]
PRODUCT APPROVED  [date] [who]

EM GO / NO-GO: GO
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If either one is REJECTED, it goes back to the engineers and that step starts over. Anything P0 (data loss, auth bypass, billing error) blocks the release no matter what the schedule says.&lt;/p&gt;

&lt;h2&gt;
  
  
  The free toolkit
&lt;/h2&gt;

&lt;p&gt;I pulled the role definitions and release gate out of my own setup into a small MIT-licensed repo: an orchestrator, an engineering manager that never writes code, backend and frontend engineers, QA, a security reviewer, an architect, and the release playbook above. There's also a &lt;code&gt;LEARNINGS.md&lt;/code&gt; with short versions of what's in this post.&lt;/p&gt;

&lt;p&gt;To be clear, it's not a new framework. If you want a full process with lots of roles, &lt;a href="https://github.com/bmad-code-org/BMAD-METHOD" rel="noopener noreferrer"&gt;BMAD-METHOD&lt;/a&gt; is much bigger and probably a better fit. This is for people who want a few opinionated roles and hard release gates. It's v0.1, it came out of my own setup, and I haven't run it end to end anywhere else, so expect rough edges. The role files are prompts, not guarantees. They don't replace reading what the agents produce.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/ageebgee-solutions/claude-agent-sdlc-toolkit" rel="noopener noreferrer"&gt;ageebgee-solutions/claude-agent-sdlc-toolkit&lt;/a&gt;. Issues and PRs are welcome, especially if you've been burned in a way I haven't listed.&lt;/p&gt;

&lt;p&gt;What have your agents gotten wrong?&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I'm Bryan, founder of &lt;a href="https://ageebgeesolutions.com/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog-distribution&amp;amp;utm_content=ai-agent-failure-modes-and-guardrails" rel="noopener noreferrer"&gt;AgeeBgee Solutions&lt;/a&gt;. We build Azure Marketplace apps.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>productivity</category>
      <category>claude</category>
    </item>
  </channel>
</rss>
