<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Robert Adamson</title>
    <description>The latest articles on DEV Community by Robert Adamson (@robertadam987_).</description>
    <link>https://dev.to/robertadam987_</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4081282%2F4198aa6c-4a89-40f3-826c-1bde258fd306.png</url>
      <title>DEV Community: Robert Adamson</title>
      <link>https://dev.to/robertadam987_</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/robertadam987_"/>
    <language>en</language>
    <item>
      <title>Most AI Agents Stop at the Answer. Xenition.com Is Built to Finish the Workflow</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Wed, 07 Oct 2026 07:06:31 +0000</pubDate>
      <link>https://dev.to/robertadam987_/most-ai-agents-stop-at-the-answer-xenitioncom-is-built-to-finish-the-workflow-4085</link>
      <guid>https://dev.to/robertadam987_/most-ai-agents-stop-at-the-answer-xenitioncom-is-built-to-finish-the-workflow-4085</guid>
      <description>&lt;p&gt;Most AI tools are very good at one thing:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Giving you an answer.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Which customers haven’t paid yet?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The AI gives you a list.&lt;/p&gt;

&lt;p&gt;You ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“What should I tell them?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The AI drafts the message.&lt;/p&gt;

&lt;p&gt;You ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Who should I follow up with first?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The AI prioritizes them.&lt;/p&gt;

&lt;p&gt;Useful?&lt;/p&gt;

&lt;p&gt;Absolutely.&lt;/p&gt;

&lt;p&gt;But then you still have to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;open the CRM&lt;/li&gt;
&lt;li&gt;find the customer&lt;/li&gt;
&lt;li&gt;open Gmail&lt;/li&gt;
&lt;li&gt;send the message&lt;/li&gt;
&lt;li&gt;update the record&lt;/li&gt;
&lt;li&gt;create a follow-up task&lt;/li&gt;
&lt;li&gt;notify your team&lt;/li&gt;
&lt;li&gt;remember to check again tomorrow&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At that point, the AI helped you think.&lt;/p&gt;

&lt;p&gt;But &lt;strong&gt;you still finished the workflow yourself.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is the problem I wanted to solve with &lt;strong&gt;Xenition.com&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Answer Is Only the Beginning
&lt;/h2&gt;

&lt;p&gt;For a long time, the basic AI workflow looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You ask
   ↓
AI answers
   ↓
You copy
   ↓
Open another app
   ↓
Paste
   ↓
Take action manually
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is still useful.&lt;/p&gt;

&lt;p&gt;But it creates a strange situation.&lt;/p&gt;

&lt;p&gt;The AI might know exactly what needs to happen next.&lt;/p&gt;

&lt;p&gt;It might know:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;which customer needs a reply&lt;/li&gt;
&lt;li&gt;which GitHub issue should be updated&lt;/li&gt;
&lt;li&gt;what should be added to Jira&lt;/li&gt;
&lt;li&gt;what message should be sent to Slack&lt;/li&gt;
&lt;li&gt;what document should be created&lt;/li&gt;
&lt;li&gt;what follow-up should happen tomorrow&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But it stops at:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Here’s what you should do.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The remaining work is still yours.&lt;/p&gt;




&lt;h1&gt;
  
  
  AI Agents Should Move From Answers to Actions
&lt;/h1&gt;

&lt;p&gt;A more useful agent should be able to move through the workflow.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User:
"Find overdue invoices and follow up with the account owners."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A chatbot might respond:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I found 8 overdue invoices.

Here are the owners and suggested messages.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An agent workflow could look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Find overdue invoices
        ↓
Identify account owners
        ↓
Prepare follow-up messages
        ↓
Request approval where needed
        ↓
Send messages
        ↓
Update records
        ↓
Log what happened
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a very different experience.&lt;/p&gt;

&lt;p&gt;The AI is no longer just helping you decide.&lt;/p&gt;

&lt;p&gt;It is helping you &lt;strong&gt;complete the task&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  This Is Why Connected Tools Matter
&lt;/h1&gt;

&lt;p&gt;An agent without tools can reason.&lt;/p&gt;

&lt;p&gt;But it cannot do much with the result.&lt;/p&gt;

&lt;p&gt;Give it access to real connected services, and suddenly the possibilities change.&lt;/p&gt;

&lt;p&gt;An agent may be able to work with things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;email&lt;/li&gt;
&lt;li&gt;calendars&lt;/li&gt;
&lt;li&gt;GitHub&lt;/li&gt;
&lt;li&gt;Slack&lt;/li&gt;
&lt;li&gt;Notion&lt;/li&gt;
&lt;li&gt;Jira&lt;/li&gt;
&lt;li&gt;CRMs&lt;/li&gt;
&lt;li&gt;cloud storage&lt;/li&gt;
&lt;li&gt;spreadsheets&lt;/li&gt;
&lt;li&gt;databases&lt;/li&gt;
&lt;li&gt;internal APIs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now the workflow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Understand the goal
      ↓
Choose the right tool
      ↓
Read the required data
      ↓
Take action
      ↓
Return the result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is much closer to how useful automation should feel.&lt;/p&gt;




&lt;h1&gt;
  
  
  This Is the Idea Behind Xenition.com
&lt;/h1&gt;

&lt;p&gt;With &lt;strong&gt;Xenition.com&lt;/strong&gt;, the goal is not to make chat the final destination.&lt;/p&gt;

&lt;p&gt;Chat is the starting point.&lt;/p&gt;

&lt;p&gt;You should be able to say what you want done, then let the system work across the tools required to complete it.&lt;/p&gt;

&lt;p&gt;The same capability can be useful in different ways.&lt;/p&gt;

&lt;h3&gt;
  
  
  Directly in chat
&lt;/h3&gt;

&lt;p&gt;You ask for something once.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Check my open issues and summarize what needs attention."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Through an agent
&lt;/h3&gt;

&lt;p&gt;The agent handles a broader goal.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Review the project and prepare today's engineering update."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Through automation
&lt;/h3&gt;

&lt;p&gt;The same workflow can run repeatedly.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Every morning
      ↓
Check project activity
      ↓
Find blockers
      ↓
Prepare update
      ↓
Send or request approval
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is where agents start becoming much more useful.&lt;/p&gt;




&lt;h1&gt;
  
  
  Agents and Automation Are Not the Same Thing
&lt;/h1&gt;

&lt;p&gt;I think this distinction matters.&lt;/p&gt;

&lt;p&gt;An &lt;strong&gt;agent&lt;/strong&gt; helps decide:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What should happen next?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;An &lt;strong&gt;automation&lt;/strong&gt; defines:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When should this workflow happen again?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;h3&gt;
  
  
  Agent
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Look at the latest support tickets and decide which ones need engineering."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent needs reasoning.&lt;/p&gt;

&lt;p&gt;It has to inspect the tickets and decide.&lt;/p&gt;

&lt;h3&gt;
  
  
  Automation
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Run this every weekday at 9 AM.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is recurrence.&lt;/p&gt;

&lt;p&gt;Put them together:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Trigger
   ↓
Agent reasons
   ↓
Uses connected tools
   ↓
Takes approved actions
   ↓
Records the result
   ↓
Runs again later
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That combination is much more powerful than chat alone.&lt;/p&gt;




&lt;h1&gt;
  
  
  But More Power Creates a New Problem
&lt;/h1&gt;

&lt;p&gt;The moment agents can actually do things, safety becomes much more important.&lt;/p&gt;

&lt;p&gt;If an AI can only write text, the main risk is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;It gives you a bad answer.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If an AI can use real tools, the risk becomes:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;It performs the wrong action.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a major difference.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Clean up old customer records."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent could interpret that as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;archive them&lt;/li&gt;
&lt;li&gt;merge them&lt;/li&gt;
&lt;li&gt;rename them&lt;/li&gt;
&lt;li&gt;delete them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are not equivalent.&lt;/p&gt;

&lt;p&gt;So agents should not have unlimited authority just because they have access to tools.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Agent Should Propose. The System Should Decide.
&lt;/h1&gt;

&lt;p&gt;This is one of the principles we have been thinking about while building Xenition.&lt;/p&gt;

&lt;p&gt;The model can decide:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What action makes sense?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But the model should not always be the final authority on:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Is this action allowed?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A better flow looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent proposes action
        ↓
System evaluates it
        ↓
Policy checks permissions
        ↓
Approval if required
        ↓
Action executes
        ↓
Result is recorded
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That separation matters.&lt;/p&gt;

&lt;p&gt;Because:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Intelligence and authority are not the same thing.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Not Every Action Needs Approval
&lt;/h1&gt;

&lt;p&gt;If every single tool call requires confirmation, automation becomes annoying very quickly.&lt;/p&gt;

&lt;p&gt;Imagine approving:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Read file?
Approve.

Read second file?
Approve.

Check calendar?
Approve.

Fetch issue?
Approve.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nobody wants that.&lt;/p&gt;

&lt;p&gt;Approval should happen when the impact changes.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Read data
→ allowed

Create draft
→ allowed

Send externally
→ approval

Delete data
→ approval

Change permissions
→ approval

Deploy
→ approval
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That creates a much better balance between usefulness and control.&lt;/p&gt;




&lt;h1&gt;
  
  
  The User Should Know What the Agent Is About to Do
&lt;/h1&gt;

&lt;p&gt;An approval screen should not say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Approve action?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is too vague.&lt;/p&gt;

&lt;p&gt;It should say something closer to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Send this message to the engineering team?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Delete these 12 records?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Publish this document externally?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The user should understand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what will happen&lt;/li&gt;
&lt;li&gt;where it will happen&lt;/li&gt;
&lt;li&gt;what will change&lt;/li&gt;
&lt;li&gt;whether it can be reversed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Otherwise approval becomes another blind click.&lt;/p&gt;




&lt;h1&gt;
  
  
  Automations Need Predictability
&lt;/h1&gt;

&lt;p&gt;There is another challenge.&lt;/p&gt;

&lt;p&gt;If an automation runs every day, you do not want the system to reinterpret the business rule differently each time.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Every morning, find high-priority customer issues and notify engineering."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The workflow should not mean one thing today and something very different next week because the model interpreted the instruction differently.&lt;/p&gt;

&lt;p&gt;That is why repeatable automation needs more structure than a normal chat prompt.&lt;/p&gt;

&lt;p&gt;You want:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;clear triggers&lt;/li&gt;
&lt;li&gt;defined tools&lt;/li&gt;
&lt;li&gt;controlled permissions&lt;/li&gt;
&lt;li&gt;known outputs&lt;/li&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;failure handling&lt;/li&gt;
&lt;li&gt;logs&lt;/li&gt;
&lt;li&gt;consistent behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AI can still reason inside the workflow.&lt;/p&gt;

&lt;p&gt;But the workflow itself should not be completely unpredictable.&lt;/p&gt;




&lt;h1&gt;
  
  
  Automation Also Needs State
&lt;/h1&gt;

&lt;p&gt;Imagine an agent runs every morning.&lt;/p&gt;

&lt;p&gt;Yesterday it already contacted Customer A.&lt;/p&gt;

&lt;p&gt;Today Customer A is still overdue.&lt;/p&gt;

&lt;p&gt;Should it send the same message again?&lt;/p&gt;

&lt;p&gt;Maybe.&lt;/p&gt;

&lt;p&gt;Maybe not.&lt;/p&gt;

&lt;p&gt;That requires state.&lt;/p&gt;

&lt;p&gt;The system needs to know things like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Already processed?
Last action?
Last result?
Waiting for reply?
Failed previously?
Needs retry?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without state, “automation” quickly becomes repeated execution instead of an actual workflow.&lt;/p&gt;




&lt;h1&gt;
  
  
  Logging Is Not Optional
&lt;/h1&gt;

&lt;p&gt;Once agents take real actions, you also need to know what happened.&lt;/p&gt;

&lt;p&gt;Not just:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Workflow completed successfully.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You need something closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;09:01 — Read 14 customer records
09:02 — Found 3 overdue invoices
09:02 — Prepared 3 messages
09:03 — 2 approved
09:04 — Messages sent
09:04 — CRM updated
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That makes the system easier to trust.&lt;/p&gt;

&lt;p&gt;It also makes failures easier to understand.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Useful Agent Workflow Has Several Layers
&lt;/h1&gt;

&lt;p&gt;The more I work on this problem, the more I think useful agent systems need more than just a smart model.&lt;/p&gt;

&lt;p&gt;They need something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Intent
  ↓
Reasoning
  ↓
Tools
  ↓
Permissions
  ↓
Approval
  ↓
Execution
  ↓
State
  ↓
Audit trail
  ↓
Automation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model is important.&lt;/p&gt;

&lt;p&gt;But it is only one part of the system.&lt;/p&gt;




&lt;h1&gt;
  
  
  Why I’m Building Xenition This Way
&lt;/h1&gt;

&lt;p&gt;This is one of the ideas behind &lt;strong&gt;Xenition.com&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I do not want AI to only answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Here is what you should do.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I want it to help close the distance between:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“I want this done.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“It’s done.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That means giving agents access to useful tools.&lt;/p&gt;

&lt;p&gt;But it also means adding:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;permission boundaries&lt;/li&gt;
&lt;li&gt;approval gates&lt;/li&gt;
&lt;li&gt;automation&lt;/li&gt;
&lt;li&gt;workflow state&lt;/li&gt;
&lt;li&gt;action history&lt;/li&gt;
&lt;li&gt;predictable execution&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because giving an agent more intelligence is not enough.&lt;/p&gt;

&lt;p&gt;It also needs an environment where it can safely use that intelligence.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Real Value Starts After the Chat Ends
&lt;/h1&gt;

&lt;p&gt;AI chat is already useful.&lt;/p&gt;

&lt;p&gt;But the most interesting part starts when the answer can become a real workflow.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Ask
↓
Answer
↓
Manual work
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;we can move toward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Intent
↓
Agent
↓
Tools
↓
Approval when needed
↓
Action
↓
Result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And when that workflow becomes repeatable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Trigger
↓
Agent
↓
Action
↓
Result
↓
Repeat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is where AI starts feeling less like a chatbot and more like actual infrastructure.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;Most AI systems are getting better at answering questions.&lt;/p&gt;

&lt;p&gt;I think the bigger opportunity is what happens &lt;strong&gt;after the answer&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Can the system:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;understand the goal?&lt;/li&gt;
&lt;li&gt;use the right tools?&lt;/li&gt;
&lt;li&gt;complete the workflow?&lt;/li&gt;
&lt;li&gt;know when it needs permission?&lt;/li&gt;
&lt;li&gt;remember what already happened?&lt;/li&gt;
&lt;li&gt;run again automatically?&lt;/li&gt;
&lt;li&gt;show exactly what it changed?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is the direction I am exploring with &lt;strong&gt;Xenition.com&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Because the best AI agent should not leave you with:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Here’s what you should do next.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It should help you get to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Done.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>automation</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Your GitHub Actions Workflow Has More Power Than Most Developers Realize</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Tue, 06 Oct 2026 04:03:37 +0000</pubDate>
      <link>https://dev.to/robertadam987_/your-github-actions-workflow-has-more-power-than-most-developers-realize-3l94</link>
      <guid>https://dev.to/robertadam987_/your-github-actions-workflow-has-more-power-than-most-developers-realize-3l94</guid>
      <description>&lt;p&gt;Most developers think of GitHub Actions as automation.&lt;/p&gt;

&lt;p&gt;Build the app.&lt;/p&gt;

&lt;p&gt;Run tests.&lt;/p&gt;

&lt;p&gt;Deploy.&lt;/p&gt;

&lt;p&gt;Publish a package.&lt;/p&gt;

&lt;p&gt;Maybe send a notification.&lt;/p&gt;

&lt;p&gt;But a workflow is not just a script runner.&lt;/p&gt;

&lt;p&gt;Depending on how you configure it, a GitHub Actions job may be able to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;read your repository&lt;/li&gt;
&lt;li&gt;modify code&lt;/li&gt;
&lt;li&gt;create releases&lt;/li&gt;
&lt;li&gt;publish packages&lt;/li&gt;
&lt;li&gt;access secrets&lt;/li&gt;
&lt;li&gt;talk to cloud providers&lt;/li&gt;
&lt;li&gt;deploy production&lt;/li&gt;
&lt;li&gt;comment on pull requests&lt;/li&gt;
&lt;li&gt;change repository state&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That means your CI pipeline is not just automation.&lt;/p&gt;

&lt;p&gt;It is part of your &lt;strong&gt;security boundary&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And many workflows have more authority than the developers maintaining them realize.&lt;/p&gt;




&lt;h2&gt;
  
  
  Start With &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Every GitHub Actions job receives a &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;GitHub creates it automatically for the job, and its permissions determine what the workflow can do inside the repository.&lt;/p&gt;

&lt;p&gt;That makes this small section of YAML extremely important:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;permissions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;contents&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;read&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without thinking about permissions explicitly, developers can easily give a workflow more authority than it actually needs.&lt;/p&gt;

&lt;p&gt;The safer idea is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Give each workflow the minimum permissions required to complete its job.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If a job only needs to check out code and run tests, it probably does not need write access.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Test Job Should Not Be Able to Publish Anything
&lt;/h1&gt;

&lt;p&gt;Consider a basic CI workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;CI&lt;/span&gt;

&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pull_request&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;

    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v6&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm ci&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ask yourself:&lt;/p&gt;

&lt;p&gt;What does this job actually need?&lt;/p&gt;

&lt;p&gt;Probably:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;permissions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;contents&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;read&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is it.&lt;/p&gt;

&lt;p&gt;It does not need permission to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;create releases&lt;/li&gt;
&lt;li&gt;modify issues&lt;/li&gt;
&lt;li&gt;publish packages&lt;/li&gt;
&lt;li&gt;write repository contents&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A useful mental model is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Permissions belong to the job, not to your general trust in GitHub Actions.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  The Most Dangerous Workflow Trigger May Surprise You
&lt;/h1&gt;

&lt;p&gt;One trigger developers should understand very carefully is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;pull_request_target&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It looks similar to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;pull_request&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But the security model is very different.&lt;/p&gt;

&lt;p&gt;A normal &lt;code&gt;pull_request&lt;/code&gt; from a fork gets strong restrictions: GitHub withholds repository secrets and gives the workflow a read-only token.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pull_request_target&lt;/code&gt;, on the other hand, runs with the trust of the base repository and can receive repository secrets and a read/write &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That is useful for things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;labeling pull requests&lt;/li&gt;
&lt;li&gt;triage&lt;/li&gt;
&lt;li&gt;authenticated checks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But it becomes dangerous if you combine it with untrusted code.&lt;/p&gt;




&lt;h1&gt;
  
  
  The “Pwn Request” Pattern
&lt;/h1&gt;

&lt;p&gt;This is the dangerous shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pull_request_target&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;

    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v6&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;ref&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ github.event.pull_request.head.sha }}&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm install&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At first glance, this looks reasonable.&lt;/p&gt;

&lt;p&gt;You want to test the contributor's code.&lt;/p&gt;

&lt;p&gt;But now you have:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Untrusted pull request code running inside a privileged workflow.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That code could modify:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;package.json&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;install scripts&lt;/li&gt;
&lt;li&gt;test scripts&lt;/li&gt;
&lt;li&gt;build scripts&lt;/li&gt;
&lt;li&gt;Makefiles&lt;/li&gt;
&lt;li&gt;configuration files&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And GitHub explicitly warns against executing fork code in a privileged &lt;code&gt;pull_request_target&lt;/code&gt; workflow.&lt;/p&gt;

&lt;p&gt;The problem is not checkout itself.&lt;/p&gt;

&lt;p&gt;The problem starts when you &lt;strong&gt;execute what you checked out&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  &lt;code&gt;npm install&lt;/code&gt; Is Code Execution
&lt;/h1&gt;

&lt;p&gt;This is another thing developers sometimes forget.&lt;/p&gt;

&lt;p&gt;When you run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you are not necessarily just downloading files.&lt;/p&gt;

&lt;p&gt;A dependency may have install scripts.&lt;/p&gt;

&lt;p&gt;Your own repository may have lifecycle scripts.&lt;/p&gt;

&lt;p&gt;Build tools may execute configuration.&lt;/p&gt;

&lt;p&gt;So if untrusted code can modify dependency or build configuration, running:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;inside a privileged job may effectively mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Execute code controlled by the pull request.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The same idea applies beyond npm:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;make
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;./gradlew build
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;cargo build
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker build
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CI commands are execution boundaries.&lt;/p&gt;

&lt;p&gt;Treat them that way.&lt;/p&gt;




&lt;h1&gt;
  
  
  Secrets Change the Risk Completely
&lt;/h1&gt;

&lt;p&gt;Imagine your workflow contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;API_KEY&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.API_KEY }}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now anything executing in that job may potentially interact with that secret.&lt;/p&gt;

&lt;p&gt;GitHub recommends using least-privilege credentials because jobs and actions can access secrets available to the workflow.&lt;/p&gt;

&lt;p&gt;A useful question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If this step were compromised, what could it steal or modify?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question should influence how you structure jobs.&lt;/p&gt;




&lt;h1&gt;
  
  
  Separate Untrusted Code From Privileged Operations
&lt;/h1&gt;

&lt;p&gt;Suppose your pipeline needs to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;test a pull request&lt;/li&gt;
&lt;li&gt;publish something afterward&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Do not automatically put everything in one giant privileged job.&lt;/p&gt;

&lt;p&gt;Think in trust boundaries.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Pull request code
      ↓
Build + test
No secrets
Read-only token
      ↓
Artifact
      ↓
Trusted workflow
      ↓
Publish / deploy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first job handles untrusted code.&lt;/p&gt;

&lt;p&gt;The later job handles privileged operations.&lt;/p&gt;

&lt;p&gt;That separation makes the system much easier to reason about.&lt;/p&gt;




&lt;h1&gt;
  
  
  Third-Party Actions Are Dependencies Too
&lt;/h1&gt;

&lt;p&gt;Developers carefully review npm packages.&lt;/p&gt;

&lt;p&gt;Then write:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;some-user/some-action@v2&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;without thinking twice.&lt;/p&gt;

&lt;p&gt;But third-party actions execute inside your workflow.&lt;/p&gt;

&lt;p&gt;A compromised action could potentially access repository secrets or use the workflow's &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;. GitHub recommends pinning actions to a full commit SHA when you need an immutable reference.&lt;/p&gt;

&lt;p&gt;Instead of relying only on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;vendor/action@v2&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;a more controlled setup can pin the exact commit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;vendor/action@3c2f...full-commit-sha&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Tags are convenient.&lt;/p&gt;

&lt;p&gt;But tags can move.&lt;/p&gt;

&lt;p&gt;A full commit SHA is immutable.&lt;/p&gt;




&lt;h1&gt;
  
  
  Your Workflow Has a Dependency Tree
&lt;/h1&gt;

&lt;p&gt;Most developers think their dependency tree is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;application
├── react
├── express
└── postgres client
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But there is another one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CI pipeline
├── actions/checkout
├── setup-node
├── deployment action
├── security scanner
└── custom third-party actions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those dependencies deserve review too.&lt;/p&gt;

&lt;p&gt;Because they may run with more privilege than your application code.&lt;/p&gt;




&lt;h1&gt;
  
  
  Long-Lived Cloud Credentials Are Often Unnecessary
&lt;/h1&gt;

&lt;p&gt;A common deployment setup looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;AWS_ACCESS_KEY_ID&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.AWS_ACCESS_KEY_ID }}&lt;/span&gt;
  &lt;span class="na"&gt;AWS_SECRET_ACCESS_KEY&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.AWS_SECRET_ACCESS_KEY }}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now you have long-lived credentials stored in GitHub.&lt;/p&gt;

&lt;p&gt;A better model, where supported, is short-lived authentication.&lt;/p&gt;

&lt;p&gt;GitHub Actions supports OpenID Connect so workflows can request short-lived identity tokens and exchange them with cloud providers instead of storing permanent cloud credentials.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GitHub workflow
      ↓
OIDC identity
      ↓
Cloud provider verifies workflow
      ↓
Temporary credentials
      ↓
Deploy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now there may be no permanent cloud secret sitting in the repository configuration.&lt;/p&gt;

&lt;p&gt;That is a major improvement.&lt;/p&gt;




&lt;h1&gt;
  
  
  Self-Hosted Runners Change the Threat Model
&lt;/h1&gt;

&lt;p&gt;GitHub-hosted runners are usually temporary.&lt;/p&gt;

&lt;p&gt;Self-hosted runners may not be.&lt;/p&gt;

&lt;p&gt;If your runner also has access to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;internal networks&lt;/li&gt;
&lt;li&gt;Docker sockets&lt;/li&gt;
&lt;li&gt;production databases&lt;/li&gt;
&lt;li&gt;cloud metadata&lt;/li&gt;
&lt;li&gt;SSH keys&lt;/li&gt;
&lt;li&gt;mounted files&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;then executing untrusted code there can become much more dangerous.&lt;/p&gt;

&lt;p&gt;GitHub specifically warns that compromised runners can expose secrets and other resources available to the runner environment.&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What can this machine reach that GitHub itself cannot?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That includes network access, not just stored credentials.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Green Build Does Not Mean a Safe Workflow
&lt;/h1&gt;

&lt;p&gt;Developers often see:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;✓ Build passed
✓ Tests passed
✓ Deployment passed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and assume everything is fine.&lt;/p&gt;

&lt;p&gt;But workflow security is not about whether the commands succeeded.&lt;/p&gt;

&lt;p&gt;It is about:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What authority did those commands have while they were running?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A perfectly successful workflow can still be badly designed.&lt;/p&gt;




&lt;h1&gt;
  
  
  Audit These 8 Things in Your GitHub Actions Today
&lt;/h1&gt;

&lt;h2&gt;
  
  
  1. Token Permissions
&lt;/h2&gt;

&lt;p&gt;Search your workflows for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;permissions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If it is missing, ask whether you should define it explicitly.&lt;/p&gt;

&lt;p&gt;Prefer the minimum required permission.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. &lt;code&gt;pull_request_target&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Search:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;pull_request_target&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you find it, understand exactly why it exists.&lt;/p&gt;

&lt;p&gt;Be especially careful if that workflow checks out and executes pull request code. GitHub is actively tightening policy around this event in public repositories because of its security implications.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Secrets
&lt;/h2&gt;

&lt;p&gt;Search for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;secrets.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For every secret, ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does this job really need it?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Do not make secrets available simply because a later step might use them.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Third-Party Actions
&lt;/h2&gt;

&lt;p&gt;Review every:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;who maintains this?&lt;/li&gt;
&lt;li&gt;do we still need it?&lt;/li&gt;
&lt;li&gt;is it pinned?&lt;/li&gt;
&lt;li&gt;what permissions does the job have?&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  5. Publishing
&lt;/h2&gt;

&lt;p&gt;Search for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;npm publish
docker push
gh release
terraform apply
kubectl
aws
gcloud
az
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Anything that changes an external system deserves extra scrutiny.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Untrusted Input
&lt;/h2&gt;

&lt;p&gt;Be careful when inserting values from issues, pull requests, branch names, commit messages, or other external sources into shell commands.&lt;/p&gt;

&lt;p&gt;The workflow file is code.&lt;/p&gt;

&lt;p&gt;Untrusted strings should be treated like untrusted input anywhere else.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Self-Hosted Runners
&lt;/h2&gt;

&lt;p&gt;Ask what the runner can access.&lt;/p&gt;

&lt;p&gt;A runner with internal network access has a very different blast radius from an isolated temporary runner.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Build and Deploy Separation
&lt;/h2&gt;

&lt;p&gt;If possible, separate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;build
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;deploy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your build job should not automatically inherit production authority just because deployment happens later.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Safer Mental Model
&lt;/h1&gt;

&lt;p&gt;Do not think:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;GitHub Actions runs my CI.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Think:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;GitHub Actions runs code with an identity, permissions, credentials, network access, and external capabilities.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That changes how you review a workflow.&lt;/p&gt;

&lt;p&gt;A workflow becomes much closer to a small production service than a simple YAML file.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;Most developers would never give every application endpoint administrator access.&lt;/p&gt;

&lt;p&gt;But we sometimes give CI pipelines broad permissions because:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“It is just our build workflow.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is not.&lt;/p&gt;

&lt;p&gt;Your workflow may have the ability to modify repositories, access secrets, publish packages, deploy infrastructure, and authenticate to production systems.&lt;/p&gt;

&lt;p&gt;So the next time you open:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.github/workflows/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;do not only ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Does this pipeline work?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What could this pipeline do if one step became untrusted?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question may be far more important.&lt;/p&gt;

</description>
      <category>discuss</category>
      <category>github</category>
      <category>productivity</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>We Gave AI Agents Real Tools — Then Realized “Just Ask Before Acting” Wasn’t Enough</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Mon, 05 Oct 2026 04:50:38 +0000</pubDate>
      <link>https://dev.to/robertadam987_/we-gave-ai-agents-real-tools-then-realized-just-ask-before-acting-wasnt-enough-19e3</link>
      <guid>https://dev.to/robertadam987_/we-gave-ai-agents-real-tools-then-realized-just-ask-before-acting-wasnt-enough-19e3</guid>
      <description>&lt;p&gt;Giving an AI agent tools feels like the moment it becomes truly useful.&lt;/p&gt;

&lt;p&gt;Now it can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;read files&lt;/li&gt;
&lt;li&gt;send messages&lt;/li&gt;
&lt;li&gt;call APIs&lt;/li&gt;
&lt;li&gt;update records&lt;/li&gt;
&lt;li&gt;trigger workflows&lt;/li&gt;
&lt;li&gt;create documents&lt;/li&gt;
&lt;li&gt;modify connected apps&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is when it stops being just a chatbot.&lt;/p&gt;

&lt;p&gt;It starts becoming software that can &lt;strong&gt;change things&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And that is also when the risk changes.&lt;/p&gt;

&lt;p&gt;At first, one rule sounds reasonable:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Ask the user before doing anything important.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Simple.&lt;/p&gt;

&lt;p&gt;Human-friendly.&lt;/p&gt;

&lt;p&gt;Easy to add to the prompt.&lt;/p&gt;

&lt;p&gt;But once an agent has real tools, that is not enough.&lt;/p&gt;

&lt;p&gt;Because now the model is being asked to decide two things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;What action should happen?&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Whether that action is important enough to require approval.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is too much authority to put inside the same reasoning loop.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem Starts With a Simple Workflow
&lt;/h2&gt;

&lt;p&gt;Imagine an agent connected to a few business tools.&lt;/p&gt;

&lt;p&gt;A user says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Clean up these customer records and notify the team.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The agent might decide to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;read CRM records&lt;/li&gt;
&lt;li&gt;modify customer fields&lt;/li&gt;
&lt;li&gt;merge duplicates&lt;/li&gt;
&lt;li&gt;delete old entries&lt;/li&gt;
&lt;li&gt;send a team message&lt;/li&gt;
&lt;li&gt;update a spreadsheet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Some of those actions are harmless.&lt;/p&gt;

&lt;p&gt;Some are reversible.&lt;/p&gt;

&lt;p&gt;Some are not.&lt;/p&gt;

&lt;p&gt;Now imagine the only safety rule is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Ask before anything important.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What exactly counts as important?&lt;/p&gt;

&lt;p&gt;The model has to decide.&lt;/p&gt;

&lt;p&gt;That is where things become uncomfortable.&lt;/p&gt;




&lt;h1&gt;
  
  
  “Important” Is Too Ambiguous
&lt;/h1&gt;

&lt;p&gt;To a user:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Deleting 500 records is obviously important.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To the agent:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Removing duplicates may look like a normal cleanup step.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To a developer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Sending data to an external system may be the risky part.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To security:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Accessing the data at all may require approval.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The word “important” does not define a reliable boundary.&lt;/p&gt;

&lt;p&gt;It creates interpretation.&lt;/p&gt;

&lt;p&gt;And interpretation is exactly what we should avoid for high-impact actions.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Model Should Not Decide Its Own Authority
&lt;/h1&gt;

&lt;p&gt;This became the key lesson.&lt;/p&gt;

&lt;p&gt;The agent can decide:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What should I do next?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But it should not be the final authority on:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Am I allowed to do it?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are separate responsibilities.&lt;/p&gt;

&lt;p&gt;A safer architecture looks more like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User request
    ↓
Agent proposes action
    ↓
System classifies action
    ↓
Policy checks permission
    ↓
Human approval if required
    ↓
Action executes
    ↓
Action is logged
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The approval decision happens outside the model.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Read, Write, and Destructive Are Not the Same
&lt;/h1&gt;

&lt;p&gt;One simple thing that helps is classifying actions by impact.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;h2&gt;
  
  
  Read
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;fetch records&lt;/li&gt;
&lt;li&gt;inspect documents&lt;/li&gt;
&lt;li&gt;search files&lt;/li&gt;
&lt;li&gt;read calendar data&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Write
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;update a field&lt;/li&gt;
&lt;li&gt;create a document&lt;/li&gt;
&lt;li&gt;send a message&lt;/li&gt;
&lt;li&gt;add an event&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Destructive / High Impact
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;delete data&lt;/li&gt;
&lt;li&gt;revoke access&lt;/li&gt;
&lt;li&gt;publish externally&lt;/li&gt;
&lt;li&gt;deploy&lt;/li&gt;
&lt;li&gt;move money&lt;/li&gt;
&lt;li&gt;change permissions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now the system can enforce something concrete.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;READ → allowed

WRITE → allowed or approval depending on context

DESTRUCTIVE → approval required
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is much stronger than:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Please ask before doing anything risky.”&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Prompts Are Guidance. Policies Are Boundaries.
&lt;/h1&gt;

&lt;p&gt;This distinction matters.&lt;/p&gt;

&lt;p&gt;A prompt can say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Never delete data without asking.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is useful.&lt;/p&gt;

&lt;p&gt;But prompts can be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;misunderstood&lt;/li&gt;
&lt;li&gt;forgotten&lt;/li&gt;
&lt;li&gt;overridden by context&lt;/li&gt;
&lt;li&gt;interpreted differently&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A policy layer should be deterministic.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;delete_record()
→ blocked
→ approval required
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent does not get to decide whether deletion is “important enough.”&lt;/p&gt;

&lt;p&gt;The system already knows.&lt;/p&gt;




&lt;h1&gt;
  
  
  Why This Matters More as Agents Get Better
&lt;/h1&gt;

&lt;p&gt;A weak agent often fails because it cannot complete the task.&lt;/p&gt;

&lt;p&gt;A strong agent creates a different problem:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;It can complete the task in ways you did not anticipate.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the real shift.&lt;/p&gt;

&lt;p&gt;The better the agent becomes at planning and using tools, the more important hard boundaries become.&lt;/p&gt;

&lt;p&gt;Because capability is increasing.&lt;/p&gt;

&lt;p&gt;Authority should not increase automatically with it.&lt;/p&gt;




&lt;h1&gt;
  
  
  Helpful Agents Can Still Cross a Line
&lt;/h1&gt;

&lt;p&gt;This is important.&lt;/p&gt;

&lt;p&gt;The dangerous behavior does not need to be malicious.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Organize this workspace.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The agent decides to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;archive old files&lt;/li&gt;
&lt;li&gt;move folders&lt;/li&gt;
&lt;li&gt;rename documents&lt;/li&gt;
&lt;li&gt;remove duplicates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every step may look helpful.&lt;/p&gt;

&lt;p&gt;But maybe one folder was legally required to remain unchanged.&lt;/p&gt;

&lt;p&gt;Maybe one document belonged to another team.&lt;/p&gt;

&lt;p&gt;Maybe the “duplicate” was actually a historical copy.&lt;/p&gt;

&lt;p&gt;The agent was trying to help.&lt;/p&gt;

&lt;p&gt;That does not make the action safe.&lt;/p&gt;




&lt;h1&gt;
  
  
  Human Approval Should Happen at the Right Moment
&lt;/h1&gt;

&lt;p&gt;Approval should not mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Confirm every tool call.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That would be terrible UX.&lt;/p&gt;

&lt;p&gt;The goal is to insert approval when the action crosses a meaningful boundary.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Reading data
→ no approval

Creating a draft
→ no approval

Sending externally
→ approval

Deleting
→ approval

Changing permissions
→ approval

Deploying
→ approval
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This keeps the agent useful without making it unrestricted.&lt;/p&gt;




&lt;h1&gt;
  
  
  Approval Should Explain the Action
&lt;/h1&gt;

&lt;p&gt;Another important lesson:&lt;/p&gt;

&lt;p&gt;Do not show the user:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Approve action?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is too vague.&lt;/p&gt;

&lt;p&gt;Show:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Send this message to the engineering channel?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Delete 42 archived records?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Publish this document externally?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The user should know exactly what they are approving.&lt;/p&gt;

&lt;p&gt;That means the approval layer needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;action name&lt;/li&gt;
&lt;li&gt;target&lt;/li&gt;
&lt;li&gt;scope&lt;/li&gt;
&lt;li&gt;consequence&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Not just a yes/no button.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Agent Should Propose, Not Hide
&lt;/h1&gt;

&lt;p&gt;A good pattern is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Agent proposes → system explains → human approves → tool runs&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Agent runs → explains afterward&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That difference matters a lot.&lt;/p&gt;

&lt;p&gt;Once the action already happened, approval is no longer approval.&lt;/p&gt;

&lt;p&gt;It is just notification.&lt;/p&gt;




&lt;h1&gt;
  
  
  This Became a Real Problem While Building Xenition
&lt;/h1&gt;

&lt;p&gt;We ran into this problem directly while building &lt;strong&gt;Xenition&lt;/strong&gt; at &lt;strong&gt;xenition.com&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Xenition is designed around AI agents that can work across real tools and connected services, not just generate text inside a chat box.&lt;/p&gt;

&lt;p&gt;That means an agent may need to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;read data&lt;/li&gt;
&lt;li&gt;create content&lt;/li&gt;
&lt;li&gt;update records&lt;/li&gt;
&lt;li&gt;trigger workflows&lt;/li&gt;
&lt;li&gt;interact with connected applications&lt;/li&gt;
&lt;li&gt;produce real outputs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Once agents can actually act, the permission model becomes just as important as the model itself.&lt;/p&gt;

&lt;p&gt;The early idea sounds simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Let the agent decide when it should ask for approval.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But that still puts too much responsibility inside the model.&lt;/p&gt;

&lt;p&gt;So the safer direction is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent proposes the action
        ↓
The system evaluates the action
        ↓
High-impact actions require approval
        ↓
The action executes
        ↓
The result is recorded
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the kind of boundary we are building around agent workflows in &lt;strong&gt;Xenition — &lt;a href="https://xenition.com/" rel="noopener noreferrer"&gt;https://xenition.com/&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The lesson was bigger than one product:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The model can decide what action makes sense.&lt;br&gt;&lt;br&gt;
The system should decide whether that action is allowed.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That separation is what starts turning an agent demo into something you can actually trust with real tools.&lt;/p&gt;




&lt;h1&gt;
  
  
  Audit Trails Matter Too
&lt;/h1&gt;

&lt;p&gt;Approval solves only part of the problem.&lt;/p&gt;

&lt;p&gt;You also need to know what happened later.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what tool was called&lt;/li&gt;
&lt;li&gt;what data changed&lt;/li&gt;
&lt;li&gt;when it happened&lt;/li&gt;
&lt;li&gt;who approved it&lt;/li&gt;
&lt;li&gt;what the agent requested&lt;/li&gt;
&lt;li&gt;what the final result was&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is why agent systems need an action ledger or audit trail.&lt;/p&gt;

&lt;p&gt;If something goes wrong, “the agent did something” is not enough.&lt;/p&gt;

&lt;p&gt;You need evidence.&lt;/p&gt;




&lt;h1&gt;
  
  
  Task-Scoped Permissions Are Even Better
&lt;/h1&gt;

&lt;p&gt;There is another improvement I think agent systems need.&lt;/p&gt;

&lt;p&gt;Do not give the agent every permission it may ever need.&lt;/p&gt;

&lt;p&gt;Give it what the &lt;strong&gt;current task&lt;/strong&gt; needs.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;h3&gt;
  
  
  Task: Summarize customer feedback
&lt;/h3&gt;

&lt;p&gt;Needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;read support tickets&lt;/li&gt;
&lt;li&gt;read CRM notes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Does not need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;delete customer&lt;/li&gt;
&lt;li&gt;modify billing&lt;/li&gt;
&lt;li&gt;publish anything&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Task: Prepare a campaign draft
&lt;/h3&gt;

&lt;p&gt;Needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;read campaign data&lt;/li&gt;
&lt;li&gt;create draft&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Does not need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;publish campaign&lt;/li&gt;
&lt;li&gt;charge customers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same agent.&lt;/p&gt;

&lt;p&gt;Different task.&lt;/p&gt;

&lt;p&gt;Different authority.&lt;/p&gt;

&lt;p&gt;That reduces blast radius dramatically.&lt;/p&gt;




&lt;h1&gt;
  
  
  Fail Closed
&lt;/h1&gt;

&lt;p&gt;One more rule:&lt;/p&gt;

&lt;p&gt;If the permission system fails, the action should stop.&lt;/p&gt;

&lt;p&gt;Bad:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;policy error
→ continue
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Better:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;policy error
→ block
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This sounds obvious.&lt;/p&gt;

&lt;p&gt;But guardrails that fail open are not really guardrails.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Simple Model That Works
&lt;/h1&gt;

&lt;p&gt;For each action, ask:&lt;/p&gt;

&lt;h2&gt;
  
  
  What is it?
&lt;/h2&gt;

&lt;p&gt;Read, write, destructive?&lt;/p&gt;

&lt;h2&gt;
  
  
  What does it affect?
&lt;/h2&gt;

&lt;p&gt;One file? One customer? Production?&lt;/p&gt;

&lt;h2&gt;
  
  
  Is it reversible?
&lt;/h2&gt;

&lt;p&gt;Can we undo it easily?&lt;/p&gt;

&lt;h2&gt;
  
  
  Does it leave the system?
&lt;/h2&gt;

&lt;p&gt;Is data being sent externally?&lt;/p&gt;

&lt;h2&gt;
  
  
  Does it need approval?
&lt;/h2&gt;

&lt;p&gt;If yes, stop before execution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is the result logged?
&lt;/h2&gt;

&lt;p&gt;Can we reconstruct what happened?&lt;/p&gt;

&lt;p&gt;That simple framework catches a surprising amount.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Bigger Lesson
&lt;/h1&gt;

&lt;p&gt;When agents only generated text, safety mostly meant:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don’t say the wrong thing.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now that agents can use real tools, safety increasingly means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don’t do the wrong thing.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That requires more than prompt engineering.&lt;/p&gt;

&lt;p&gt;It requires:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;permissions&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;policy enforcement&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;approval gates&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;task-scoped access&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;audit logs&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;fail-closed behavior&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The model can decide what action makes sense.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The system should decide whether that action is allowed.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;Giving AI agents real tools is what makes them powerful.&lt;/p&gt;

&lt;p&gt;It is also what makes them dangerous if the authority model is vague.&lt;/p&gt;

&lt;p&gt;“Ask before acting” sounds safe.&lt;/p&gt;

&lt;p&gt;But it still asks the model to decide when it needs permission.&lt;/p&gt;

&lt;p&gt;That is the wrong place to put the boundary.&lt;/p&gt;

&lt;p&gt;The safer model is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Let the agent propose.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Let policy decide.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Let the human approve when the impact is high.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Because the moment an AI agent can change the real world, permission stops being a prompt.&lt;/p&gt;

&lt;p&gt;It becomes part of the architecture.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>programming</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>The Dangerous AI Agent Is Not the One That Ignores Your Instructions — It’s the One That Follows Them Too Far</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Sun, 04 Oct 2026 04:27:02 +0000</pubDate>
      <link>https://dev.to/robertadam987_/the-dangerous-ai-agent-is-not-the-one-that-ignores-your-instructions-its-the-one-that-follows-266p</link>
      <guid>https://dev.to/robertadam987_/the-dangerous-ai-agent-is-not-the-one-that-ignores-your-instructions-its-the-one-that-follows-266p</guid>
      <description>&lt;p&gt;We often think the dangerous AI agent is the one that refuses instructions.&lt;/p&gt;

&lt;p&gt;The one that goes rogue.&lt;/p&gt;

&lt;p&gt;The one that ignores what we asked.&lt;/p&gt;

&lt;p&gt;But there is another failure mode that may be more realistic:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The agent understands the goal perfectly — and pursues it too aggressively.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a much harder problem.&lt;/p&gt;

&lt;p&gt;Because the agent may not be “disobeying” you at all.&lt;/p&gt;

&lt;p&gt;It may simply be optimizing for the objective without understanding where its authority should stop.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Goal Can Be Correct While the Action Is Wrong
&lt;/h2&gt;

&lt;p&gt;Imagine you ask an agent:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Find why the deployment failed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The goal is reasonable.&lt;/p&gt;

&lt;p&gt;The agent starts investigating.&lt;/p&gt;

&lt;p&gt;It reads logs.&lt;/p&gt;

&lt;p&gt;Checks config.&lt;/p&gt;

&lt;p&gt;Inspects CI.&lt;/p&gt;

&lt;p&gt;Queries cloud resources.&lt;/p&gt;

&lt;p&gt;Looks at credentials.&lt;/p&gt;

&lt;p&gt;Calls internal services.&lt;/p&gt;

&lt;p&gt;Maybe even changes something to test a theory.&lt;/p&gt;

&lt;p&gt;At each step, the agent may believe:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;This helps me complete the task.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And that is exactly the problem.&lt;/p&gt;

&lt;p&gt;The question is not only:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Does the agent understand the goal?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is also:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Does the agent understand what it is allowed to do while pursuing that goal?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are two different things.&lt;/p&gt;




&lt;h1&gt;
  
  
  Goal Alignment Is Not Permission Alignment
&lt;/h1&gt;

&lt;p&gt;This distinction matters.&lt;/p&gt;

&lt;p&gt;A goal says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What should be achieved?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Permissions say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What actions are allowed?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="1hpl1y"&lt;br&gt;
Goal:&lt;br&gt;
Fix the production outage.&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


That does not automatically mean:



```text id="ihdveu"
Permission:
Restart services
Change firewall rules
Rotate credentials
Modify database records
Deploy code
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;But if the agent has access to those capabilities, it may decide they are useful.&lt;/p&gt;

&lt;p&gt;The agent can be perfectly aligned with the task and still cross a boundary.&lt;/p&gt;


&lt;h1&gt;
  
  
  Prompts Are Not Security Boundaries
&lt;/h1&gt;

&lt;p&gt;A common pattern is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Do not touch production.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Do not delete anything.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Ask before deploying.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those instructions are useful.&lt;/p&gt;

&lt;p&gt;But they are not strong security controls.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because they depend on the agent interpreting and remembering the rule correctly.&lt;/p&gt;

&lt;p&gt;A stronger system makes forbidden actions technically unavailable.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="4i41wi"&lt;br&gt;
Please do not access production.&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


prefer:



```text id="8mc8ra"
production_credentials = unavailable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="ev4xsx"&lt;br&gt;
Do not call external services.&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


prefer:



```text id="8n79xt"
network_access = allowlist only
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="1pzd88"&lt;br&gt;
Ask before deployment.&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


prefer:



```text id="70g7fh"
deploy = human approval required
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That is a much safer model.&lt;/p&gt;


&lt;h1&gt;
  
  
  Capability Is Not Authority
&lt;/h1&gt;

&lt;p&gt;An agent may technically be capable of doing something.&lt;/p&gt;

&lt;p&gt;That does not mean the current task should authorize it.&lt;/p&gt;

&lt;p&gt;This is one of the biggest design mistakes I see in agent workflows.&lt;/p&gt;

&lt;p&gt;A coding agent may have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;shell access&lt;/li&gt;
&lt;li&gt;Git access&lt;/li&gt;
&lt;li&gt;cloud credentials&lt;/li&gt;
&lt;li&gt;package manager access&lt;/li&gt;
&lt;li&gt;database access&lt;/li&gt;
&lt;li&gt;deployment tools&lt;/li&gt;
&lt;li&gt;network access&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But if the task is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Fix a button alignment bug.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Why should it inherit all of that?&lt;/p&gt;

&lt;p&gt;The better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What does this task actually require?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;


&lt;h1&gt;
  
  
  Permissions Should Be Task-Scoped
&lt;/h1&gt;

&lt;p&gt;Imagine two tasks.&lt;/p&gt;
&lt;h3&gt;
  
  
  Task A — Fix CSS
&lt;/h3&gt;

&lt;p&gt;The agent probably needs:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="kq1v5k"&lt;br&gt;
read frontend files&lt;br&gt;
write frontend files&lt;br&gt;
run frontend tests&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


It probably does not need:



```text id="b9wsj1"
cloud admin access
production database access
npm publish
deployment credentials
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h3&gt;
  
  
  Task B — Prepare a Release
&lt;/h3&gt;

&lt;p&gt;Now the agent may need:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="fx39n5"&lt;br&gt;
build&lt;br&gt;
test&lt;br&gt;
create release artifact&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


But publishing could still require:



```text id="s0h91l"
human approval
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Same agent.&lt;/p&gt;

&lt;p&gt;Different task.&lt;/p&gt;

&lt;p&gt;Different authority.&lt;/p&gt;

&lt;p&gt;That feels like the safer model.&lt;/p&gt;


&lt;h1&gt;
  
  
  Helpful Agents Can Still Be Dangerous
&lt;/h1&gt;

&lt;p&gt;This is the uncomfortable part.&lt;/p&gt;

&lt;p&gt;The agent does not need malicious intent.&lt;/p&gt;

&lt;p&gt;It may simply reason:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I need more information.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So it reads another file.&lt;/p&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I need to verify this.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So it calls another tool.&lt;/p&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I can fix this directly.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So it modifies something.&lt;/p&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“The fix should be deployed to confirm it.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And suddenly the agent has crossed several boundaries while still pursuing the original goal.&lt;/p&gt;

&lt;p&gt;Every step may look locally reasonable.&lt;/p&gt;

&lt;p&gt;The full sequence may not be.&lt;/p&gt;


&lt;h1&gt;
  
  
  This Is Similar to Architecture Drift
&lt;/h1&gt;

&lt;p&gt;A single action may look harmless.&lt;/p&gt;

&lt;p&gt;But a chain of individually reasonable actions can create a bad outcome.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="ljz6dp"&lt;br&gt;
Read logs&lt;br&gt;
↓&lt;br&gt;
Inspect credentials&lt;br&gt;
↓&lt;br&gt;
Query internal API&lt;br&gt;
↓&lt;br&gt;
Modify config&lt;br&gt;
↓&lt;br&gt;
Restart service&lt;br&gt;
↓&lt;br&gt;
Deploy change&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


Maybe no individual step looked outrageous.

But the agent gradually expanded its own scope.

That is why task boundaries need to exist outside the model.

---

# Human Approval Should Be About Escalation

Human approval is most useful when the agent is about to increase its authority.

For example:

Require approval before:

- modifying production
- deleting files
- installing new dependencies
- publishing packages
- accessing secrets
- changing permissions
- sending data externally
- deploying
- touching infrastructure

The agent can still move quickly.

But high-impact actions create a checkpoint.

---

# Default to Read-Only

A very practical rule:

&amp;gt; **Start agents read-only whenever possible.**

Let them:

- inspect
- analyze
- propose
- explain
- generate plans

Then promote permissions only when necessary.

For example:



```text id="q9m37c"
Stage 1:
read only

Stage 2:
write project files

Stage 3:
run approved commands

Stage 4:
sensitive action requires human approval
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That creates a natural escalation path.&lt;/p&gt;


&lt;h1&gt;
  
  
  Make Permission Changes Visible
&lt;/h1&gt;

&lt;p&gt;If the agent needs more access, it should say so explicitly.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I can continue analyzing with current permissions.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;To complete this step, I need write access to &lt;code&gt;config/&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Deployment requires production credentials and approval.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That makes authority visible.&lt;/p&gt;

&lt;p&gt;Silent escalation is the dangerous part.&lt;/p&gt;


&lt;h1&gt;
  
  
  Network Access Matters Too
&lt;/h1&gt;

&lt;p&gt;Developers often think only about credentials.&lt;/p&gt;

&lt;p&gt;But network position matters as well.&lt;/p&gt;

&lt;p&gt;An agent running inside your machine may have access to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;VPN routes&lt;/li&gt;
&lt;li&gt;internal DNS&lt;/li&gt;
&lt;li&gt;localhost services&lt;/li&gt;
&lt;li&gt;company APIs&lt;/li&gt;
&lt;li&gt;metadata endpoints&lt;/li&gt;
&lt;li&gt;unauthenticated internal tools&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Even without credentials, it may still reach things that the public internet cannot.&lt;/p&gt;

&lt;p&gt;So sandboxing should include:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;filesystem&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;credentials&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;tools&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;and:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;network egress&lt;/strong&gt;&lt;/p&gt;


&lt;h1&gt;
  
  
  Fail Closed, Not Open
&lt;/h1&gt;

&lt;p&gt;Suppose a policy hook fails.&lt;/p&gt;

&lt;p&gt;What happens?&lt;/p&gt;

&lt;p&gt;Bad design:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="2z7i12"&lt;br&gt;
policy check fails&lt;br&gt;
↓&lt;br&gt;
agent continues&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


Better:



```text id="36p98h"
policy check fails
↓
action blocked
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Security boundaries should fail closed.&lt;/p&gt;

&lt;p&gt;If the system cannot determine whether an action is allowed, the safest default is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do not perform it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;


&lt;h1&gt;
  
  
  Log What the Agent Actually Did
&lt;/h1&gt;

&lt;p&gt;Permissions tell you what an agent &lt;strong&gt;could&lt;/strong&gt; do.&lt;/p&gt;

&lt;p&gt;Logs tell you what it &lt;strong&gt;did&lt;/strong&gt; do.&lt;/p&gt;

&lt;p&gt;For meaningful agent workflows, I want an audit trail containing things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tool calls&lt;/li&gt;
&lt;li&gt;commands&lt;/li&gt;
&lt;li&gt;file writes&lt;/li&gt;
&lt;li&gt;network requests&lt;/li&gt;
&lt;li&gt;approval requests&lt;/li&gt;
&lt;li&gt;permission escalations&lt;/li&gt;
&lt;li&gt;deployment actions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Not just:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Task completed successfully.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The summary is not enough.&lt;/p&gt;

&lt;p&gt;The actions matter.&lt;/p&gt;


&lt;h1&gt;
  
  
  Separate the Goal From the Policy
&lt;/h1&gt;

&lt;p&gt;One useful architecture is to keep them independent.&lt;/p&gt;
&lt;h3&gt;
  
  
  Agent
&lt;/h3&gt;

&lt;p&gt;Figures out:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What should I do next?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Policy layer
&lt;/h3&gt;

&lt;p&gt;Checks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is this action allowed?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The agent should not be the final authority on both.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="zfko91"&lt;br&gt;
Agent:&lt;br&gt;
"Run production migration."&lt;/p&gt;

&lt;p&gt;Policy:&lt;br&gt;
"Production writes require human approval."&lt;/p&gt;

&lt;p&gt;Result:&lt;br&gt;
Blocked pending approval.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


That is much stronger than telling the agent:

&amp;gt; “Remember to ask first.”

---

# A Simple Agent Permission Model

For each task, define:

## Read

What can the agent inspect?

## Write

What can it modify?

## Execute

Which commands can it run?

## Network

Which destinations can it reach?

## Credentials

Which identities can it use?

## Escalation

Which actions require approval?

That is already enough to make agent workflows much easier to reason about.

---

# Before Giving an Agent a Task, Ask These Questions

### What is the goal?

Be specific.

### What is the minimum authority needed?

Do not inherit everything by default.

### What actions should require approval?

Define them before execution.

### What should be impossible?

Enforce that technically.

### What happens if the agent misunderstands the boundary?

The system should still remain safe.

### Can I reconstruct what happened later?

Keep an audit trail.

---

# The Bigger Lesson

We spend a lot of time trying to make agents understand our goals better.

That is important.

But understanding the goal is only half of the problem.

The other half is:

&amp;gt; **Understanding authority.**

An agent might know exactly what you want.

It might even find a very effective way to achieve it.

And that way may still be unacceptable.

---

# Final Thought

The dangerous AI agent is not always the one that says:

&amp;gt; **“I won’t follow your instructions.”**

Sometimes it is the one that says:

&amp;gt; **“I understand exactly what you want. I’ll do whatever is necessary to achieve it.”**

That is why production agent systems need more than good prompts.

They need:

**permissions**

**boundaries**

**approval gates**

**sandboxing**

**network controls**

**audit trails**

because:

&amp;gt; **A goal tells the agent what success looks like.**

&amp;gt; **Authority tells it how far it is allowed to go.**

And those two things should never be confused.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>The More Context You Give Your AI Coding Agent, the Worse It Can Get</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Sat, 03 Oct 2026 03:52:59 +0000</pubDate>
      <link>https://dev.to/robertadam987_/the-more-context-you-give-your-ai-coding-agent-the-worse-it-can-get-4d40</link>
      <guid>https://dev.to/robertadam987_/the-more-context-you-give-your-ai-coding-agent-the-worse-it-can-get-4d40</guid>
      <description>&lt;p&gt;We keep hearing the same advice:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Give the AI more context.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Add the README.&lt;/p&gt;

&lt;p&gt;Add &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Add architecture docs.&lt;/p&gt;

&lt;p&gt;Add logs.&lt;/p&gt;

&lt;p&gt;Add previous decisions.&lt;/p&gt;

&lt;p&gt;Add the whole repository.&lt;/p&gt;

&lt;p&gt;Add memory from previous sessions.&lt;/p&gt;

&lt;p&gt;Sounds reasonable.&lt;/p&gt;

&lt;p&gt;But there is a problem:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;More context does not always mean better understanding.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sometimes it means more noise.&lt;/p&gt;

&lt;p&gt;More stale assumptions.&lt;/p&gt;

&lt;p&gt;More conflicting instructions.&lt;/p&gt;

&lt;p&gt;More irrelevant files.&lt;/p&gt;

&lt;p&gt;And more chances for the agent to focus on the wrong thing.&lt;/p&gt;

&lt;p&gt;That is the part I think developers need to pay more attention to.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Assumption: More Context = Better AI
&lt;/h2&gt;

&lt;p&gt;It makes sense at first.&lt;/p&gt;

&lt;p&gt;If the agent knows more about the codebase, it should make better decisions.&lt;/p&gt;

&lt;p&gt;Right?&lt;/p&gt;

&lt;p&gt;Sometimes.&lt;/p&gt;

&lt;p&gt;But imagine giving a developer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;400 files&lt;/li&gt;
&lt;li&gt;12 architecture documents&lt;/li&gt;
&lt;li&gt;8 old incident reports&lt;/li&gt;
&lt;li&gt;3 outdated migration plans&lt;/li&gt;
&lt;li&gt;6 instruction files&lt;/li&gt;
&lt;li&gt;40 pages of logs&lt;/li&gt;
&lt;li&gt;previous agent memory&lt;/li&gt;
&lt;li&gt;the current task&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;and then asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Fix this bug.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is not automatically helpful.&lt;/p&gt;

&lt;p&gt;That is a lot of information to sort through.&lt;/p&gt;

&lt;p&gt;The same problem can happen with AI agents.&lt;/p&gt;




&lt;h1&gt;
  
  
  Context Has Quality, Not Just Quantity
&lt;/h1&gt;

&lt;p&gt;Not all context is equally useful.&lt;/p&gt;

&lt;p&gt;Some context is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Relevant&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Some is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Outdated&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Some is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Conflicting&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Some is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wrong&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Some is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Technically correct but irrelevant to the current task&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you give all of it equal weight, the agent has to figure out what matters.&lt;/p&gt;

&lt;p&gt;And that is where mistakes start.&lt;/p&gt;




&lt;h1&gt;
  
  
  Example: The Old Architecture Doc
&lt;/h1&gt;

&lt;p&gt;Suppose your current system uses:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="ixn1f3"&lt;br&gt;
Controller&lt;br&gt;
   ↓&lt;br&gt;
Service&lt;br&gt;
   ↓&lt;br&gt;
Repository&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


But an old architecture document still says:



```text id="8ylsl1"
Controller
   ↓
Data Layer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;You ask the agent to add a feature.&lt;/p&gt;

&lt;p&gt;Now the agent has two sources of truth.&lt;/p&gt;

&lt;p&gt;Which one should it trust?&lt;/p&gt;

&lt;p&gt;Maybe it follows the code.&lt;/p&gt;

&lt;p&gt;Maybe it follows the documentation.&lt;/p&gt;

&lt;p&gt;Maybe it blends both.&lt;/p&gt;

&lt;p&gt;And now you get a new pattern that never existed before.&lt;/p&gt;

&lt;p&gt;The problem was not lack of context.&lt;/p&gt;

&lt;p&gt;The problem was &lt;strong&gt;bad context hygiene&lt;/strong&gt;.&lt;/p&gt;


&lt;h1&gt;
  
  
  Stale Context Is Worse Than Missing Context
&lt;/h1&gt;

&lt;p&gt;Missing context usually creates uncertainty.&lt;/p&gt;

&lt;p&gt;Stale context can create confidence in the wrong direction.&lt;/p&gt;

&lt;p&gt;That is more dangerous.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Three months ago:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“All payments go through Provider A.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Today:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Half the system has moved to Provider B.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But the agent still carries the old rule in memory.&lt;/p&gt;

&lt;p&gt;Now it confidently implements the wrong integration.&lt;/p&gt;

&lt;p&gt;That is why persistent memory can be useful and dangerous at the same time.&lt;/p&gt;


&lt;h1&gt;
  
  
  Conflicting Instructions Create Quiet Problems
&lt;/h1&gt;

&lt;p&gt;Imagine the agent reads:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="hp3on9"&lt;br&gt;
AGENTS.md:&lt;br&gt;
Use service classes for all business logic.&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


Then another file says:



```text id="x68ti8"
README:
Keep business logic inside route handlers.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Then an old task note says:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="14479z"&lt;br&gt;
Avoid adding new service layers.&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


All three may have been correct at different times.

Now they coexist.

The agent has to resolve the conflict.

That is not a safe default.

---

# More Tokens Do Not Mean More Attention

This is another important point.

A bigger context window gives the agent access to more information.

It does not guarantee equal attention to every piece of information.

If you include:

- hundreds of files
- long logs
- old discussions
- huge docs

the important detail may become harder to surface.

The real problem becomes:

&amp;gt; **Can the agent find the right context at the right moment?**

That is different from:

&amp;gt; **Can the agent fit everything into the prompt?**

---

# The Goal Should Not Be Maximum Context

I think the better goal is:

&amp;gt; **Minimum sufficient context.**

Give the agent enough information to make the right decision.

Not everything you have.

For example, if the task is:

&amp;gt; “Fix a validation bug in checkout.”

The agent probably needs:

- checkout flow
- validation rules
- related tests
- relevant data model
- current architecture constraints

It probably does not need:

- email service docs
- analytics history
- unrelated migration logs
- old design discussions
- every frontend component

More is not automatically better.

---

# Use Progressive Context

A better pattern is:



```text id="k7ov9i"
Start small
↓
Give relevant files
↓
Let the agent inspect
↓
Add more only when needed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="dxypv0"&lt;br&gt;
Dump everything&lt;br&gt;
↓&lt;br&gt;
Hope the agent finds what matters&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


This is basically progressive disclosure for coding agents.

Let the agent earn more context as the task requires it.

---

# Ask the Agent What It Needs

This is surprisingly useful.

Instead of giving the whole repo immediately, ask:



```text id="wd8gzg"
Before changing anything:

1. What information do you need?
2. Which files are likely relevant?
3. What assumptions are you currently making?
4. What context would reduce uncertainty?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Now the agent tells you what it is missing.&lt;/p&gt;

&lt;p&gt;That is much better than blindly adding more.&lt;/p&gt;


&lt;h1&gt;
  
  
  Separate Permanent Context From Task Context
&lt;/h1&gt;

&lt;p&gt;I think teams should split context into two categories.&lt;/p&gt;
&lt;h2&gt;
  
  
  Permanent
&lt;/h2&gt;

&lt;p&gt;Things that should almost always be true:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;coding standards&lt;/li&gt;
&lt;li&gt;architecture boundaries&lt;/li&gt;
&lt;li&gt;security rules&lt;/li&gt;
&lt;li&gt;naming conventions&lt;/li&gt;
&lt;li&gt;ownership rules&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Task-specific
&lt;/h2&gt;

&lt;p&gt;Things relevant only to the current job:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one bug report&lt;/li&gt;
&lt;li&gt;one feature requirement&lt;/li&gt;
&lt;li&gt;one set of logs&lt;/li&gt;
&lt;li&gt;one module&lt;/li&gt;
&lt;li&gt;one incident&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Mixing both into one giant blob makes reasoning harder.&lt;/p&gt;


&lt;h1&gt;
  
  
  Keep Permanent Rules Short
&lt;/h1&gt;

&lt;p&gt;This is important.&lt;/p&gt;

&lt;p&gt;Your agent instructions should not become a novel.&lt;/p&gt;

&lt;p&gt;If your &lt;code&gt;AGENTS.md&lt;/code&gt; is 5,000 lines long, developers probably do not read it carefully either.&lt;/p&gt;

&lt;p&gt;The best permanent rules are usually simple.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="w7h7wu"&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Business logic stays in services.&lt;/li&gt;
&lt;li&gt;Repositories only handle persistence.&lt;/li&gt;
&lt;li&gt;Do not add dependencies without approval.&lt;/li&gt;
&lt;li&gt;Do not modify auth rules without explicit request.&lt;/li&gt;
&lt;li&gt;Tests must cover changed behavior.
```
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Clear.&lt;/p&gt;

&lt;p&gt;Short.&lt;/p&gt;

&lt;p&gt;Hard to misinterpret.&lt;/p&gt;




&lt;h1&gt;
  
  
  Context Should Have an Expiration Date
&lt;/h1&gt;

&lt;p&gt;Some context should not live forever.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;temporary migration rules&lt;/li&gt;
&lt;li&gt;incident-specific workarounds&lt;/li&gt;
&lt;li&gt;old feature flags&lt;/li&gt;
&lt;li&gt;deprecated API behavior&lt;/li&gt;
&lt;li&gt;one-off implementation notes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the agent can remember something forever, someone needs to decide when that memory stops being valid.&lt;/p&gt;

&lt;p&gt;This is why I think agent memory needs something humans already understand:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lifecycle management.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Context should be:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;created&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;reviewed&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;updated&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;expired&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;deleted&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Just like code and documentation.&lt;/p&gt;




&lt;h1&gt;
  
  
  Make Sources Visible
&lt;/h1&gt;

&lt;p&gt;Another good habit:&lt;/p&gt;

&lt;p&gt;Do not let context appear as one anonymous blob.&lt;/p&gt;

&lt;p&gt;The agent should know where information came from.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="l6tqad"&lt;br&gt;
Source: current code&lt;br&gt;
Source: AGENTS.md&lt;br&gt;
Source: architecture decision record&lt;br&gt;
Source: incident from June&lt;br&gt;
Source: previous agent memory&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


Why?

Because source matters.

Current production code should probably outweigh a two-year-old planning document.

Without provenance, everything can look equally trustworthy.

---

# Ask the Agent to Surface Conflicts

Before implementation, try:



```text id="gprxfr"
Review the available context.

Identify:

- conflicting instructions
- outdated assumptions
- duplicated rules
- unclear sources of truth
- anything that may no longer be valid

Do not change code yet.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;This is a very useful step for large codebases.&lt;/p&gt;

&lt;p&gt;You want context conflicts visible before they become code.&lt;/p&gt;


&lt;h1&gt;
  
  
  More Context Can Increase Hallucination Too
&lt;/h1&gt;

&lt;p&gt;This sounds backwards.&lt;/p&gt;

&lt;p&gt;But imagine the agent sees 10 partial references to a system behavior.&lt;/p&gt;

&lt;p&gt;None gives the complete picture.&lt;/p&gt;

&lt;p&gt;It may combine them into a plausible explanation.&lt;/p&gt;

&lt;p&gt;That explanation can sound very confident.&lt;/p&gt;

&lt;p&gt;And still be wrong.&lt;/p&gt;

&lt;p&gt;The issue is not always missing information.&lt;/p&gt;

&lt;p&gt;Sometimes it is &lt;strong&gt;too many incomplete signals&lt;/strong&gt;.&lt;/p&gt;


&lt;h1&gt;
  
  
  Context Is Part of the Architecture Now
&lt;/h1&gt;

&lt;p&gt;We usually think architecture means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;services&lt;/li&gt;
&lt;li&gt;databases&lt;/li&gt;
&lt;li&gt;queues&lt;/li&gt;
&lt;li&gt;APIs&lt;/li&gt;
&lt;li&gt;boundaries&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But in AI-assisted development, context becomes part of the system too.&lt;/p&gt;

&lt;p&gt;Because context influences:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what the agent believes&lt;/li&gt;
&lt;li&gt;what it changes&lt;/li&gt;
&lt;li&gt;which patterns it follows&lt;/li&gt;
&lt;li&gt;which assumptions it preserves&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That means context needs engineering discipline too.&lt;/p&gt;


&lt;h1&gt;
  
  
  A Simple Context Checklist
&lt;/h1&gt;

&lt;p&gt;Before giving an AI agent more information, ask:&lt;/p&gt;
&lt;h3&gt;
  
  
  Is this relevant?
&lt;/h3&gt;

&lt;p&gt;If not, leave it out.&lt;/p&gt;
&lt;h3&gt;
  
  
  Is this still true?
&lt;/h3&gt;

&lt;p&gt;If you are not sure, verify it.&lt;/p&gt;
&lt;h3&gt;
  
  
  Does it conflict with another source?
&lt;/h3&gt;

&lt;p&gt;Resolve that first.&lt;/p&gt;
&lt;h3&gt;
  
  
  Is there a newer source?
&lt;/h3&gt;

&lt;p&gt;Prefer the newer one.&lt;/p&gt;
&lt;h3&gt;
  
  
  Does the agent need this now?
&lt;/h3&gt;

&lt;p&gt;Maybe later is better.&lt;/p&gt;
&lt;h3&gt;
  
  
  Will this context still be valid next month?
&lt;/h3&gt;

&lt;p&gt;If not, do not treat it as permanent memory.&lt;/p&gt;


&lt;h1&gt;
  
  
  My Preferred Workflow
&lt;/h1&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```text id="5axqqf"&lt;br&gt;
Load entire repo&lt;br&gt;
↓&lt;br&gt;
Load all docs&lt;br&gt;
↓&lt;br&gt;
Load memory&lt;br&gt;
↓&lt;br&gt;
Ask agent to work&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


I prefer:



```text id="d6w46d"
Define task
↓
Give core constraints
↓
Agent identifies needed context
↓
Load only relevant files
↓
Check for conflicts
↓
Implement
↓
Verify
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference is simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context becomes intentional.&lt;/strong&gt;&lt;/p&gt;




&lt;h1&gt;
  
  
  The Bigger Lesson
&lt;/h1&gt;

&lt;p&gt;AI coding agents do not just need more information.&lt;/p&gt;

&lt;p&gt;They need:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the right information&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;from the right source&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;at the right time&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is a much harder problem.&lt;/p&gt;

&lt;p&gt;But it is also where developers can add real value.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;We keep trying to make AI coding agents smarter by giving them more context.&lt;/p&gt;

&lt;p&gt;Sometimes that works.&lt;/p&gt;

&lt;p&gt;Sometimes it makes things worse.&lt;/p&gt;

&lt;p&gt;Because:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;More context can mean more noise.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More memory can mean more stale assumptions.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More instructions can mean more conflicts.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The goal should not be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Give the agent everything.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The goal should be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Give the agent exactly what it needs to make the right decision.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Because the best context window is not the biggest one.&lt;/p&gt;

&lt;p&gt;It is the one with the &lt;strong&gt;least irrelevant information and the clearest source of truth&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>tutorial</category>
      <category>discuss</category>
    </item>
    <item>
      <title>AI Helps You Code Faster. So Why Are You Still Shipping Slowly?</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Wed, 30 Sep 2026 05:50:09 +0000</pubDate>
      <link>https://dev.to/robertadam987_/ai-helps-you-code-faster-so-why-are-you-still-shipping-slowly-dl1</link>
      <guid>https://dev.to/robertadam987_/ai-helps-you-code-faster-so-why-are-you-still-shipping-slowly-dl1</guid>
      <description>&lt;p&gt;You generated a feature in 20 minutes.&lt;/p&gt;

&lt;p&gt;Three days later, it still hasn’t reached production.&lt;/p&gt;

&lt;p&gt;The code exists. But the tests are failing, the review is waiting, and someone just noticed that the feature solves the wrong problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What actually got faster?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Writing the first version.&lt;/p&gt;

&lt;p&gt;That matters. But users don’t use first drafts. They use working software.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coding Is Only One Part of Shipping
&lt;/h2&gt;

&lt;p&gt;Consider a simple feature: adding a discount code at checkout.&lt;/p&gt;

&lt;p&gt;AI can quickly generate the input, validation logic, and API endpoint.&lt;/p&gt;

&lt;p&gt;But you still need answers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can customers combine discounts?&lt;/li&gt;
&lt;li&gt;What happens when a code expires during checkout?&lt;/li&gt;
&lt;li&gt;Does the discount apply before or after tax?&lt;/li&gt;
&lt;li&gt;Can someone reuse a single-use code?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A feature can look finished while these questions remain unanswered.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI can turn assumptions into code faster than you can notice them.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Faster Generation Can Create a Bigger Queue
&lt;/h2&gt;

&lt;p&gt;Imagine a team producing twice as many changes while its review capacity stays the same.&lt;/p&gt;

&lt;p&gt;More code arrives. More code waits.&lt;/p&gt;

&lt;p&gt;Large changes make this harder. A reviewer must understand the intent, inspect the implementation, and check how it affects the rest of the system.&lt;/p&gt;

&lt;p&gt;Generating another feature doesn’t clear that queue.&lt;/p&gt;

&lt;p&gt;Sometimes, the most useful next step is helping an existing change reach production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Passing Tests Doesn’t Mean You Solved the Right Problem
&lt;/h2&gt;

&lt;p&gt;Suppose you ask AI to implement discounts, then ask it to write tests.&lt;/p&gt;

&lt;p&gt;It might assume that discounts can be combined—and write tests confirming that behavior.&lt;/p&gt;

&lt;p&gt;Everything passes.&lt;/p&gt;

&lt;p&gt;The business wanted one discount per order.&lt;/p&gt;

&lt;p&gt;The problem is that &lt;strong&gt;the implementation and the tests can share the same wrong assumption.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Define the expected behavior before generating either.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five Ways to Turn Faster Coding Into Faster Delivery
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Describe “done” before asking for code
&lt;/h3&gt;

&lt;p&gt;“A discount feature” is vague.&lt;/p&gt;

&lt;p&gt;“One valid discount per order, with expired codes rejected and the updated total shown before payment” gives you something concrete to verify.&lt;/p&gt;

&lt;p&gt;Clear requirements reduce guessing and rework.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Ask for smaller changes
&lt;/h3&gt;

&lt;p&gt;Avoid asking an agent to rebuild checkout in one request.&lt;/p&gt;

&lt;p&gt;Start with one behavior. Inspect it. Verify it. Then move on.&lt;/p&gt;

&lt;p&gt;A smaller change is easier to understand, review, and fix.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Review the assumptions first
&lt;/h3&gt;

&lt;p&gt;Before examining every line, ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What decisions did the AI make?&lt;/li&gt;
&lt;li&gt;Which existing behavior changed?&lt;/li&gt;
&lt;li&gt;What happens when something fails?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Finding a wrong assumption early can save more time than polishing the code built around it.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Test the awkward cases
&lt;/h3&gt;

&lt;p&gt;The happy path is only the beginning.&lt;/p&gt;

&lt;p&gt;Try an expired code, two quick submissions, a failed request, and a customer refreshing during checkout.&lt;/p&gt;

&lt;p&gt;Choose tests based on what could hurt the user or the business.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Measure delivery, not generated code
&lt;/h3&gt;

&lt;p&gt;Lines of code and completed prompts don’t tell you whether customers received something useful.&lt;/p&gt;

&lt;p&gt;Track how long changes take to reach production, how often they need rework, and what breaks after release.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A smaller feature that works is more valuable than a larger feature waiting for approval.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Try This on Your Next Feature
&lt;/h2&gt;

&lt;p&gt;Write three acceptance criteria. Ask AI for the smallest implementation that meets them. Check its assumptions. Test one important failure case.&lt;/p&gt;

&lt;p&gt;Then follow the change all the way to production.&lt;/p&gt;

&lt;p&gt;AI can shorten the coding step. To shorten delivery, you also need to address whatever keeps the finished code waiting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where does your work get stuck most often: requirements, review, testing, or deployment?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>discuss</category>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
    </item>
    <item>
      <title>AI Can Make Every Local Decision Look Reasonable — While Making the System Worse</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Tue, 29 Sep 2026 03:46:12 +0000</pubDate>
      <link>https://dev.to/robertadam987_/ai-can-make-every-local-decision-look-reasonable-while-making-the-system-worse-11p0</link>
      <guid>https://dev.to/robertadam987_/ai-can-make-every-local-decision-look-reasonable-while-making-the-system-worse-11p0</guid>
      <description>&lt;p&gt;AI coding agents are very good at making small decisions.&lt;/p&gt;

&lt;p&gt;Move this logic into a helper.&lt;/p&gt;

&lt;p&gt;Add a cache here.&lt;/p&gt;

&lt;p&gt;Create a service for that.&lt;/p&gt;

&lt;p&gt;Split this module.&lt;/p&gt;

&lt;p&gt;Add another abstraction.&lt;/p&gt;

&lt;p&gt;Introduce one more dependency.&lt;/p&gt;

&lt;p&gt;Each decision can look perfectly reasonable.&lt;/p&gt;

&lt;p&gt;That is exactly the problem.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A system can get worse even when every local decision looks correct.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is not really an AI problem.&lt;/p&gt;

&lt;p&gt;It is an architecture problem that AI can now accelerate.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Dangerous Part Is That Nothing Looks Obviously Wrong
&lt;/h2&gt;

&lt;p&gt;Imagine this over a few weeks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Task 1 → Add validation
Task 2 → Extract helper
Task 3 → Add cache
Task 4 → Add service
Task 5 → Add retry logic
Task 6 → Refactor module
Task 7 → Add another integration
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every pull request passes.&lt;/p&gt;

&lt;p&gt;Every change looks clean.&lt;/p&gt;

&lt;p&gt;Every test is green.&lt;/p&gt;

&lt;p&gt;But six weeks later:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;two services now own the same business rule&lt;/li&gt;
&lt;li&gt;validation exists in three places&lt;/li&gt;
&lt;li&gt;caching is inconsistent&lt;/li&gt;
&lt;li&gt;dependencies point in both directions&lt;/li&gt;
&lt;li&gt;nobody knows which layer owns what&lt;/li&gt;
&lt;li&gt;a “simple” change touches eight files&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No single commit destroyed the architecture.&lt;/p&gt;

&lt;p&gt;The system drifted there.&lt;/p&gt;




&lt;h1&gt;
  
  
  Local Correctness Is Not System Correctness
&lt;/h1&gt;

&lt;p&gt;This is the key idea.&lt;/p&gt;

&lt;p&gt;An AI agent usually works on the task in front of it.&lt;/p&gt;

&lt;p&gt;You ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Add retry support.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It finds the easiest reasonable place to add retries.&lt;/p&gt;

&lt;p&gt;You ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Reuse this logic.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It extracts a helper.&lt;/p&gt;

&lt;p&gt;You ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Clean up this module.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It introduces a new abstraction.&lt;/p&gt;

&lt;p&gt;Each answer may be correct for that prompt.&lt;/p&gt;

&lt;p&gt;But architecture is not just a collection of locally correct choices.&lt;/p&gt;

&lt;p&gt;Architecture is about how those choices interact over time.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Agent Optimizes for the Task
&lt;/h1&gt;

&lt;p&gt;The developer has a different responsibility.&lt;/p&gt;

&lt;p&gt;The agent asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How do I complete this task successfully?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The developer should ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What does this change do to the whole system?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are not the same question.&lt;/p&gt;

&lt;p&gt;The agent may optimize for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;fewer lines&lt;/li&gt;
&lt;li&gt;cleaner separation&lt;/li&gt;
&lt;li&gt;less duplication&lt;/li&gt;
&lt;li&gt;faster implementation&lt;/li&gt;
&lt;li&gt;easier reuse&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The system may actually need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;stronger boundaries&lt;/li&gt;
&lt;li&gt;fewer abstractions&lt;/li&gt;
&lt;li&gt;one clear owner&lt;/li&gt;
&lt;li&gt;deliberate duplication&lt;/li&gt;
&lt;li&gt;less coupling&lt;/li&gt;
&lt;li&gt;simpler dependencies&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A locally elegant change can still be globally harmful.&lt;/p&gt;




&lt;h1&gt;
  
  
  Example: The Helpful Helper
&lt;/h1&gt;

&lt;p&gt;Suppose you have this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;calculateInvoiceTotal&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One module uses it.&lt;/p&gt;

&lt;p&gt;Then another feature needs similar behavior.&lt;/p&gt;

&lt;p&gt;AI suggests:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Move it into shared/utils.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Makes sense.&lt;/p&gt;

&lt;p&gt;Later another module imports it.&lt;/p&gt;

&lt;p&gt;Then another.&lt;/p&gt;

&lt;p&gt;Then someone adds customer-specific logic.&lt;/p&gt;

&lt;p&gt;Then tax rules.&lt;/p&gt;

&lt;p&gt;Then discount behavior.&lt;/p&gt;

&lt;p&gt;Now your “shared helper” contains business logic used across the entire system.&lt;/p&gt;

&lt;p&gt;Nothing looked unreasonable at the time.&lt;/p&gt;

&lt;p&gt;But the architecture quietly changed.&lt;/p&gt;

&lt;p&gt;What started as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Invoice owns invoice logic
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Everyone depends on shared/utils
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is architecture drift.&lt;/p&gt;




&lt;h1&gt;
  
  
  AI Makes Architecture Drift Faster
&lt;/h1&gt;

&lt;p&gt;Before AI, introducing five new abstractions took effort.&lt;/p&gt;

&lt;p&gt;Now it can happen in minutes.&lt;/p&gt;

&lt;p&gt;That changes the economics.&lt;/p&gt;

&lt;p&gt;Developers can generate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;new services&lt;/li&gt;
&lt;li&gt;new interfaces&lt;/li&gt;
&lt;li&gt;new layers&lt;/li&gt;
&lt;li&gt;new adapters&lt;/li&gt;
&lt;li&gt;new helpers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;faster than they can evaluate whether those things should exist.&lt;/p&gt;

&lt;p&gt;The risk is not that AI generates obviously terrible architecture.&lt;/p&gt;

&lt;p&gt;The more interesting risk is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;It generates plausible architecture faster than teams can maintain a coherent design.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Green Tests Do Not Protect Architecture
&lt;/h1&gt;

&lt;p&gt;This is important.&lt;/p&gt;

&lt;p&gt;Your test suite may say:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Everything works.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That does not mean:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The system is getting better.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Tests can catch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;broken behavior&lt;/li&gt;
&lt;li&gt;regressions&lt;/li&gt;
&lt;li&gt;invalid outputs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They usually do not catch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;unclear ownership&lt;/li&gt;
&lt;li&gt;unnecessary abstractions&lt;/li&gt;
&lt;li&gt;dependency creep&lt;/li&gt;
&lt;li&gt;duplicated business rules&lt;/li&gt;
&lt;li&gt;architectural drift&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can have:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;100% tests passing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and still have a codebase becoming harder to change every week.&lt;/p&gt;




&lt;h1&gt;
  
  
  Code Review Can Miss It Too
&lt;/h1&gt;

&lt;p&gt;A reviewer sees one pull request.&lt;/p&gt;

&lt;p&gt;Maybe 80 lines.&lt;/p&gt;

&lt;p&gt;The diff looks reasonable.&lt;/p&gt;

&lt;p&gt;But architecture problems often become visible only across many changes.&lt;/p&gt;

&lt;p&gt;PR 1 adds a helper.&lt;/p&gt;

&lt;p&gt;PR 2 adds a service.&lt;/p&gt;

&lt;p&gt;PR 3 adds another dependency.&lt;/p&gt;

&lt;p&gt;PR 4 reuses the helper.&lt;/p&gt;

&lt;p&gt;PR 5 bypasses the service.&lt;/p&gt;

&lt;p&gt;Individually:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fine.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Collectively:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Messy system.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is why reviewing only the current diff is not always enough.&lt;/p&gt;




&lt;h1&gt;
  
  
  Ask One Extra Question During AI-Assisted Review
&lt;/h1&gt;

&lt;p&gt;When reviewing AI-generated code, ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If we repeat this pattern 20 times, what does the system look like?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question is surprisingly useful.&lt;/p&gt;

&lt;p&gt;If every feature adds:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;another service&lt;/li&gt;
&lt;li&gt;another config&lt;/li&gt;
&lt;li&gt;another queue&lt;/li&gt;
&lt;li&gt;another shared helper&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;then the pattern itself may be wrong.&lt;/p&gt;

&lt;p&gt;Architecture is often about what happens when a decision gets repeated.&lt;/p&gt;




&lt;h1&gt;
  
  
  Protect Ownership Boundaries
&lt;/h1&gt;

&lt;p&gt;One of the easiest ways to reduce drift is to make ownership explicit.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Payments owns payment rules.
Orders owns order lifecycle.
Users owns identity.
Notifications only sends messages.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then tell the agent:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do not move business rules across these boundaries unless explicitly requested.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is better than letting the model decide architectural ownership from scratch every time.&lt;/p&gt;




&lt;h1&gt;
  
  
  Give AI Constraints, Not Just Goals
&lt;/h1&gt;

&lt;p&gt;Bad prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Refactor this to make it cleaner.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Better:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Refactor this module.

Constraints:

- preserve current architectural boundaries
- do not introduce new dependencies
- do not create new shared abstractions
- keep business logic in the existing domain
- propose changes before implementing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not to make AI less useful.&lt;/p&gt;

&lt;p&gt;It is to stop every task from becoming a mini architecture redesign.&lt;/p&gt;




&lt;h1&gt;
  
  
  Ask for the Architecture Impact Before the Code
&lt;/h1&gt;

&lt;p&gt;Before implementation, try:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Before changing code, explain:

1. Which architectural boundary this touches.
2. Which modules will depend on the change.
3. Whether this introduces a new abstraction.
4. Whether this creates a new dependency.
5. Whether similar logic already exists.
6. How this affects future changes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That forces the discussion one level above syntax.&lt;/p&gt;




&lt;h1&gt;
  
  
  Watch for Dependency Direction
&lt;/h1&gt;

&lt;p&gt;One signal I pay attention to is dependency direction.&lt;/p&gt;

&lt;p&gt;Healthy systems usually have clear relationships.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Controller
   ↓
Service
   ↓
Repository
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then AI introduces:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Repository
   ↓
Utility
   ↓
Service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now layers start depending on each other in strange ways.&lt;/p&gt;

&lt;p&gt;Each import may compile.&lt;/p&gt;

&lt;p&gt;But the system becomes harder to reason about.&lt;/p&gt;

&lt;p&gt;If dependency direction becomes unclear, architecture is usually drifting.&lt;/p&gt;




&lt;h1&gt;
  
  
  Shared Code Is Not Always Better Code
&lt;/h1&gt;

&lt;p&gt;AI often tries to remove duplication.&lt;/p&gt;

&lt;p&gt;Usually that is good.&lt;/p&gt;

&lt;p&gt;But sometimes two pieces of code only look similar today.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;calculateSellerFee()
calculateCreatorFee()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AI may combine them into:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;calculateFee(type)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cleaner?&lt;/p&gt;

&lt;p&gt;Maybe.&lt;/p&gt;

&lt;p&gt;But if seller and creator rules evolve differently, you now have one abstraction carrying two separate concepts.&lt;/p&gt;

&lt;p&gt;A little duplication can be cheaper than the wrong abstraction.&lt;/p&gt;

&lt;p&gt;This is something humans still need to judge carefully.&lt;/p&gt;




&lt;h1&gt;
  
  
  Architecture Needs Invariants
&lt;/h1&gt;

&lt;p&gt;I think AI-heavy teams need explicit architectural rules.&lt;/p&gt;

&lt;p&gt;Not hundreds of pages.&lt;/p&gt;

&lt;p&gt;Just a few things that must remain true.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- Controllers contain no business logic.
- Domains cannot access each other's database tables directly.
- Shared utilities cannot contain business rules.
- External APIs go through adapters.
- Background jobs call services instead of repositories directly.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now both humans and agents have something concrete to check.&lt;/p&gt;

&lt;p&gt;These are architecture invariants.&lt;/p&gt;




&lt;h1&gt;
  
  
  Use AI to Detect Drift Too
&lt;/h1&gt;

&lt;p&gt;AI can also help find the problem it creates.&lt;/p&gt;

&lt;p&gt;Periodically ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Review this codebase for architecture drift.

Look for:

- duplicated responsibilities
- unclear module ownership
- circular dependencies
- new abstractions with little value
- shared utilities containing business logic
- inconsistent patterns
- dependency direction violations

Do not refactor anything yet.

Rank the findings by impact.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That can be a very useful review.&lt;/p&gt;

&lt;p&gt;The important part:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do not refactor anything yet.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;First detect.&lt;/p&gt;

&lt;p&gt;Then decide.&lt;/p&gt;




&lt;h1&gt;
  
  
  Small Changes Need System-Level Review Too
&lt;/h1&gt;

&lt;p&gt;The dangerous phrase is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“It’s only a small change.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Small changes repeated hundreds of times become architecture.&lt;/p&gt;

&lt;p&gt;Every small change teaches the codebase a pattern.&lt;/p&gt;

&lt;p&gt;If that pattern is weak, AI can replicate it extremely quickly.&lt;/p&gt;

&lt;p&gt;That is why local quality is not enough.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Simple Review Checklist
&lt;/h1&gt;

&lt;p&gt;Before merging an AI-generated change, ask:&lt;/p&gt;

&lt;h3&gt;
  
  
  Does this introduce a new abstraction?
&lt;/h3&gt;

&lt;p&gt;If yes, do we actually need it?&lt;/p&gt;

&lt;h3&gt;
  
  
  Does this move responsibility?
&lt;/h3&gt;

&lt;p&gt;If yes, is the new owner correct?&lt;/p&gt;

&lt;h3&gt;
  
  
  Does this add another dependency?
&lt;/h3&gt;

&lt;p&gt;Could an existing boundary handle it?&lt;/p&gt;

&lt;h3&gt;
  
  
  Does similar logic already exist?
&lt;/h3&gt;

&lt;p&gt;Avoid competing patterns.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens if we repeat this approach 20 times?
&lt;/h3&gt;

&lt;p&gt;This one catches a lot.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is the system easier to explain after this change?
&lt;/h3&gt;

&lt;p&gt;If not, the “cleaner code” may not actually be cleaner.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Real Risk Is Plausible Decisions at Scale
&lt;/h1&gt;

&lt;p&gt;AI does not need to make obviously bad decisions to damage architecture.&lt;/p&gt;

&lt;p&gt;It only needs to make thousands of reasonable decisions without enough system-level coordination.&lt;/p&gt;

&lt;p&gt;That is the uncomfortable part.&lt;/p&gt;

&lt;p&gt;A codebase rarely becomes difficult because someone deliberately designed it badly.&lt;/p&gt;

&lt;p&gt;It usually becomes difficult through:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;hundreds of decisions that made sense at the time.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;AI can now generate those decisions much faster.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;AI is very good at answering:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What is a reasonable solution to this task?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But software architecture requires another question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What happens to the system if we keep making decisions like this?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are different levels of reasoning.&lt;/p&gt;

&lt;p&gt;And that is why developers still need to protect the bigger picture.&lt;/p&gt;

&lt;p&gt;Because:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Every local decision can look reasonable while the system slowly gets worse.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The agent owns the task.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You still own the architecture.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>discuss</category>
      <category>ai</category>
      <category>programming</category>
      <category>security</category>
    </item>
    <item>
      <title>AI Can Fix the Bug Before You Understand It — That’s More Dangerous Than It Sounds</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Mon, 28 Sep 2026 05:01:27 +0000</pubDate>
      <link>https://dev.to/robertadam987_/ai-can-fix-the-bug-before-you-understand-it-thats-more-dangerous-than-it-sounds-466j</link>
      <guid>https://dev.to/robertadam987_/ai-can-fix-the-bug-before-you-understand-it-thats-more-dangerous-than-it-sounds-466j</guid>
      <description>&lt;p&gt;A bug appears.&lt;/p&gt;

&lt;p&gt;You paste the error into your AI coding agent.&lt;/p&gt;

&lt;p&gt;Thirty seconds later:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Fixed ✅”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You run the app.&lt;/p&gt;

&lt;p&gt;It works.&lt;/p&gt;

&lt;p&gt;Great.&lt;/p&gt;

&lt;p&gt;Except there is one problem:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;You still don’t know why it broke.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And I think this is becoming one of the most overlooked problems in AI-assisted development.&lt;/p&gt;

&lt;p&gt;AI is making bug fixing faster.&lt;/p&gt;

&lt;p&gt;But it can also make &lt;strong&gt;understanding optional&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is dangerous.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Old Debugging Loop
&lt;/h2&gt;

&lt;p&gt;Before AI, debugging usually looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bug
↓
Reproduce
↓
Read logs
↓
Form hypothesis
↓
Test hypothesis
↓
Find root cause
↓
Fix
↓
Add regression test
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It was slower.&lt;/p&gt;

&lt;p&gt;But during that process, you built a mental model of the system.&lt;/p&gt;

&lt;p&gt;You learned:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;where data flows&lt;/li&gt;
&lt;li&gt;what assumptions exist&lt;/li&gt;
&lt;li&gt;which component owns what&lt;/li&gt;
&lt;li&gt;what can fail&lt;/li&gt;
&lt;li&gt;how the system behaves under pressure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now the workflow can become:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bug
↓
Ask AI
↓
Patch appears
↓
App works
↓
Move on
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fast?&lt;/p&gt;

&lt;p&gt;Yes.&lt;/p&gt;

&lt;p&gt;Safe?&lt;/p&gt;

&lt;p&gt;Not always.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Fix Is Not the Same as Understanding
&lt;/h1&gt;

&lt;p&gt;Suppose your app crashes because a value is unexpectedly &lt;code&gt;null&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;AI adds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The crash disappears.&lt;/p&gt;

&lt;p&gt;But did we actually fix the problem?&lt;/p&gt;

&lt;p&gt;Maybe the real issue was:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the user should never have been null&lt;/li&gt;
&lt;li&gt;a database query failed&lt;/li&gt;
&lt;li&gt;an async race happened&lt;/li&gt;
&lt;li&gt;state was loaded too late&lt;/li&gt;
&lt;li&gt;an API returned invalid data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The patch removed the symptom.&lt;/p&gt;

&lt;p&gt;It may not have removed the cause.&lt;/p&gt;

&lt;p&gt;This is the important distinction:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A successful patch does not prove the diagnosis was correct.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  AI Is Very Good at Symptom Removal
&lt;/h1&gt;

&lt;p&gt;AI often sees:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;error → likely patch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That can be useful.&lt;/p&gt;

&lt;p&gt;But some “fixes” simply:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;add a null check&lt;/li&gt;
&lt;li&gt;catch an exception&lt;/li&gt;
&lt;li&gt;retry the operation&lt;/li&gt;
&lt;li&gt;increase a timeout&lt;/li&gt;
&lt;li&gt;suppress an error&lt;/li&gt;
&lt;li&gt;weaken validation&lt;/li&gt;
&lt;li&gt;add a fallback&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The application stops failing.&lt;/p&gt;

&lt;p&gt;But the deeper problem may still exist.&lt;/p&gt;

&lt;p&gt;That is how temporary fixes become permanent technical debt.&lt;/p&gt;




&lt;h1&gt;
  
  
  Ask for the Root Cause Before the Fix
&lt;/h1&gt;

&lt;p&gt;One habit helps a lot.&lt;/p&gt;

&lt;p&gt;Instead of asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Fix this bug.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Try:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Do not modify the code yet.

First explain:

1. What is failing?
2. What triggered the failure?
3. What state should have existed?
4. What state actually existed?
5. What is the most likely root cause?
6. What evidence supports that conclusion?
7. What other explanations are possible?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only then ask for the fix.&lt;/p&gt;

&lt;p&gt;That keeps you involved in the reasoning.&lt;/p&gt;




&lt;h1&gt;
  
  
  Use This Debugging Flow
&lt;/h1&gt;

&lt;p&gt;A better AI-assisted debugging workflow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Reproduce
↓
Observe
↓
Explain
↓
Form hypothesis
↓
Verify hypothesis
↓
Fix
↓
Regression test
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI can help at every step.&lt;/p&gt;

&lt;p&gt;But do not skip the middle.&lt;/p&gt;

&lt;p&gt;That middle is where understanding happens.&lt;/p&gt;




&lt;h1&gt;
  
  
  Ask the AI to Prove Its Diagnosis
&lt;/h1&gt;

&lt;p&gt;If the agent says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“The bug is caused by a race condition.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What evidence makes you think that?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What observation would prove this diagnosis wrong?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That second question is extremely useful.&lt;/p&gt;

&lt;p&gt;A good debugging process should be falsifiable.&lt;/p&gt;

&lt;p&gt;Not just confident.&lt;/p&gt;




&lt;h1&gt;
  
  
  Separate Diagnosis From Repair
&lt;/h1&gt;

&lt;p&gt;This is one of the best changes you can make.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1 — Diagnose
&lt;/h3&gt;

&lt;p&gt;Ask the AI to inspect the failure and explain it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2 — Verify
&lt;/h3&gt;

&lt;p&gt;Check logs, state, tests, timing, or data.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3 — Repair
&lt;/h3&gt;

&lt;p&gt;Only after the cause is reasonably clear.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4 — Prove
&lt;/h3&gt;

&lt;p&gt;Add a regression test that fails before the fix and passes after it.&lt;/p&gt;

&lt;p&gt;That is much stronger than:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“The error disappeared.”&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Require a Regression Test
&lt;/h1&gt;

&lt;p&gt;Every meaningful bug fix should answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What test would have caught this before production?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If the answer is “none,” the bug can easily come back.&lt;/p&gt;

&lt;p&gt;A good regression test should:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;reproduce the original failure&lt;/li&gt;
&lt;li&gt;fail before the patch&lt;/li&gt;
&lt;li&gt;pass after the patch&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Now the fix becomes part of the system’s knowledge.&lt;/p&gt;

&lt;p&gt;Not just the AI conversation.&lt;/p&gt;




&lt;h1&gt;
  
  
  Be Careful With “It Works Now”
&lt;/h1&gt;

&lt;p&gt;“It works now” is one of the weakest debugging signals.&lt;/p&gt;

&lt;p&gt;Maybe:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the failure is intermittent&lt;/li&gt;
&lt;li&gt;cached state changed&lt;/li&gt;
&lt;li&gt;timing changed&lt;/li&gt;
&lt;li&gt;the test data is different&lt;/li&gt;
&lt;li&gt;a retry succeeded&lt;/li&gt;
&lt;li&gt;the error moved somewhere else&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why does it work now?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If nobody can answer that, the debugging is not finished.&lt;/p&gt;




&lt;h1&gt;
  
  
  Your Team Needs to Own the Explanation
&lt;/h1&gt;

&lt;p&gt;This is the part that worries me most.&lt;/p&gt;

&lt;p&gt;Imagine an AI fixes bugs all day.&lt;/p&gt;

&lt;p&gt;Everything ships faster.&lt;/p&gt;

&lt;p&gt;But six months later, your team knows less about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;data flow&lt;/li&gt;
&lt;li&gt;failure modes&lt;/li&gt;
&lt;li&gt;edge cases&lt;/li&gt;
&lt;li&gt;architecture&lt;/li&gt;
&lt;li&gt;production behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The codebase may improve.&lt;/p&gt;

&lt;p&gt;Your understanding of it may not.&lt;/p&gt;

&lt;p&gt;That creates a new kind of risk.&lt;/p&gt;




&lt;h1&gt;
  
  
  I Think of It as Debugging Debt
&lt;/h1&gt;

&lt;p&gt;We already talk about technical debt.&lt;/p&gt;

&lt;p&gt;AI can create another kind:&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Debugging Debt&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;You solved the incident.&lt;/p&gt;

&lt;p&gt;But you never learned why it happened.&lt;/p&gt;

&lt;p&gt;The next failure may come from the same underlying cause in a different place.&lt;/p&gt;

&lt;p&gt;And now your team has to rediscover everything from scratch.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Best Use of AI in Debugging
&lt;/h1&gt;

&lt;p&gt;AI is incredibly useful for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;reading logs&lt;/li&gt;
&lt;li&gt;tracing code paths&lt;/li&gt;
&lt;li&gt;finding suspicious changes&lt;/li&gt;
&lt;li&gt;generating hypotheses&lt;/li&gt;
&lt;li&gt;comparing stack traces&lt;/li&gt;
&lt;li&gt;creating repro tests&lt;/li&gt;
&lt;li&gt;suggesting instrumentation&lt;/li&gt;
&lt;li&gt;identifying likely failure points&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But the final question should still be human-owned:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do we actually understand why this failed?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the difference between:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;patching&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;and&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;debugging&lt;/strong&gt;&lt;/p&gt;




&lt;h1&gt;
  
  
  A Simple Prompt I Use
&lt;/h1&gt;

&lt;p&gt;When something breaks, try this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Do not fix anything yet.

Help me debug this systematically.

1. Summarize the failure.
2. Trace the likely execution path.
3. List the top 3 root-cause hypotheses.
4. For each hypothesis, give evidence for and against it.
5. Tell me what logs, tests, or observations would confirm it.
6. Wait before proposing code changes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This forces the AI to act more like a debugging partner than a patch generator.&lt;/p&gt;




&lt;h1&gt;
  
  
  My Final Rule
&lt;/h1&gt;

&lt;p&gt;Before accepting an AI-generated bug fix, I want to be able to answer:&lt;/p&gt;

&lt;h3&gt;
  
  
  What broke?
&lt;/h3&gt;

&lt;h3&gt;
  
  
  Why did it break?
&lt;/h3&gt;

&lt;h3&gt;
  
  
  Why does this fix work?
&lt;/h3&gt;

&lt;h3&gt;
  
  
  What would prove this fix is wrong?
&lt;/h3&gt;

&lt;h3&gt;
  
  
  What regression test protects us now?
&lt;/h3&gt;

&lt;p&gt;If I cannot answer those questions, I probably accepted a patch too early.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;AI can fix bugs faster than most developers ever could manually.&lt;/p&gt;

&lt;p&gt;That is powerful.&lt;/p&gt;

&lt;p&gt;But speed creates a temptation:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Skip understanding. Accept the patch. Move on.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That works until the next incident.&lt;/p&gt;

&lt;p&gt;The goal should not be:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI fixed it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The goal should be:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;We understand why it broke, and now it cannot break the same way again.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because a bug you fixed but never understood is not really knowledge your team owns.&lt;/p&gt;

&lt;p&gt;It is knowledge you temporarily borrowed from the AI.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>learning</category>
      <category>tools</category>
    </item>
    <item>
      <title>Your AI Coding Agent Says “Tests Pass.” But Did It Actually Run Them?</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Sun, 27 Sep 2026 05:07:41 +0000</pubDate>
      <link>https://dev.to/robertadam987_/your-ai-coding-agent-says-tests-pass-but-did-it-actually-run-them-4684</link>
      <guid>https://dev.to/robertadam987_/your-ai-coding-agent-says-tests-pass-but-did-it-actually-run-them-4684</guid>
      <description>&lt;p&gt;AI coding agents are getting very good at finishing tasks.&lt;/p&gt;

&lt;p&gt;They modify files.&lt;/p&gt;

&lt;p&gt;Fix errors.&lt;/p&gt;

&lt;p&gt;Write tests.&lt;/p&gt;

&lt;p&gt;Run commands.&lt;/p&gt;

&lt;p&gt;Then they end with something like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;All tests pass ✅&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And most of us move on.&lt;/p&gt;

&lt;p&gt;But there is an important question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Did the agent actually run the tests it claims passed?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sounds obvious.&lt;/p&gt;

&lt;p&gt;It isn’t.&lt;/p&gt;

&lt;p&gt;Because in AI-assisted development, we are starting to trust &lt;strong&gt;summaries&lt;/strong&gt; instead of &lt;strong&gt;evidence&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  “Tests Pass” Is a Claim
&lt;/h2&gt;

&lt;p&gt;Imagine an agent says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Implementation complete.

✓ Tests pass
✓ Build succeeds
✓ No lint errors
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That looks reassuring.&lt;/p&gt;

&lt;p&gt;But you still do not know:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what command it ran&lt;/li&gt;
&lt;li&gt;whether it ran the full test suite&lt;/li&gt;
&lt;li&gt;whether some tests were skipped&lt;/li&gt;
&lt;li&gt;whether the command exited successfully&lt;/li&gt;
&lt;li&gt;whether the output came from the latest code&lt;/li&gt;
&lt;li&gt;whether it ran tests at all&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The message is a summary.&lt;/p&gt;

&lt;p&gt;It is not proof.&lt;/p&gt;




&lt;h2&gt;
  
  
  This Is a New Kind of Trust Problem
&lt;/h2&gt;

&lt;p&gt;Before coding agents, developers usually ran commands directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You saw the output.&lt;/p&gt;

&lt;p&gt;You saw the failures.&lt;/p&gt;

&lt;p&gt;You saw the exit code.&lt;/p&gt;

&lt;p&gt;With agents, the workflow can become:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Developer asks for feature
        ↓
Agent changes code
        ↓
Agent runs something
        ↓
Agent summarizes result
        ↓
Developer trusts summary
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is now another layer between you and the actual verification.&lt;/p&gt;

&lt;p&gt;That layer can be wrong.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Agent May Have Run Only Part of the Tests
&lt;/h2&gt;

&lt;p&gt;Suppose your project has:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;unit tests
integration tests
API tests
end-to-end tests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm run &lt;span class="nb"&gt;test&lt;/span&gt;:unit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Everything passes.&lt;/p&gt;

&lt;p&gt;Then it reports:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;All tests pass.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Technically, some tests passed.&lt;/p&gt;

&lt;p&gt;But the full application was never verified.&lt;/p&gt;

&lt;p&gt;That difference matters.&lt;/p&gt;




&lt;h2&gt;
  
  
  It May Be Reporting Stale Results
&lt;/h2&gt;

&lt;p&gt;Another easy failure mode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent runs tests
        ↓
Tests pass
        ↓
Agent changes code again
        ↓
Agent reports "tests pass"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The statement was true earlier.&lt;/p&gt;

&lt;p&gt;It may no longer be true now.&lt;/p&gt;

&lt;p&gt;Verification should happen against the &lt;strong&gt;final state of the code&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Passing Tests Can Still Mean the Wrong Tests
&lt;/h2&gt;

&lt;p&gt;There is another problem.&lt;/p&gt;

&lt;p&gt;AI writes the feature.&lt;/p&gt;

&lt;p&gt;Then AI writes the tests.&lt;/p&gt;

&lt;p&gt;Then AI runs those tests.&lt;/p&gt;

&lt;p&gt;They pass.&lt;/p&gt;

&lt;p&gt;Great.&lt;/p&gt;

&lt;p&gt;Except both the implementation and the tests may share the same misunderstanding.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Wrong requirement interpretation
        ↓
AI writes code
        ↓
AI writes tests for that interpretation
        ↓
Tests pass
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Everything is green.&lt;/p&gt;

&lt;p&gt;The feature is still wrong.&lt;/p&gt;

&lt;p&gt;So there are really two questions:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Did the tests run?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Were they the right tests?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Both matter.&lt;/p&gt;




&lt;h1&gt;
  
  
  Require Evidence, Not Confidence
&lt;/h1&gt;

&lt;p&gt;I have started preferring a simple rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If an agent makes a verification claim, ask for the evidence behind it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of accepting:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Tests pass.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ask for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Show me:

1. The exact command you ran.
2. The exit code.
3. How many tests ran.
4. How many failed.
5. How many were skipped.
6. Whether this was run after the final code change.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the claim becomes inspectable.&lt;/p&gt;




&lt;h1&gt;
  
  
  Give Your Agent a Verification Contract
&lt;/h1&gt;

&lt;p&gt;You can make this part of your normal coding-agent instructions.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Never say "tests pass" unless you actually ran the relevant test command.

When reporting verification, include:

- exact command
- exit code
- number of tests
- failures
- skipped tests
- build status
- lint/typecheck status

If you did not run something, say "not verified."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line is especially important:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If you did not verify it, say so.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;“I did not run the integration tests” is much more useful than false confidence.&lt;/p&gt;




&lt;h1&gt;
  
  
  Separate Implementation From Verification
&lt;/h1&gt;

&lt;p&gt;For important changes, do not let the same workflow both create and certify the result.&lt;/p&gt;

&lt;p&gt;A stronger pattern is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent implements feature
        ↓
Independent test command
        ↓
CI verifies
        ↓
Human reviews result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For higher-risk changes, you can go further:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent A → implementation

Agent B → adversarial review

CI → tests/build/security checks

Human → final approval
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is independence.&lt;/p&gt;

&lt;p&gt;The system that produced the code should not be your only source of confidence that the code is correct.&lt;/p&gt;




&lt;h1&gt;
  
  
  CI Should Be the Source of Truth
&lt;/h1&gt;

&lt;p&gt;The agent's message should be treated as useful context.&lt;/p&gt;

&lt;p&gt;Not final authority.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Agent says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;All tests pass.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;CI says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;3 integration tests failed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Trust CI.&lt;/p&gt;

&lt;p&gt;This is why traditional engineering systems still matter even when AI writes more of the code.&lt;/p&gt;

&lt;p&gt;You still want:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;CI&lt;/li&gt;
&lt;li&gt;deterministic tests&lt;/li&gt;
&lt;li&gt;build checks&lt;/li&gt;
&lt;li&gt;type checks&lt;/li&gt;
&lt;li&gt;linting&lt;/li&gt;
&lt;li&gt;security scans&lt;/li&gt;
&lt;li&gt;deployment gates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AI does not replace those systems.&lt;/p&gt;

&lt;p&gt;It makes them more important.&lt;/p&gt;




&lt;h1&gt;
  
  
  Be Careful With “Fixed”
&lt;/h1&gt;

&lt;p&gt;The same principle applies to other agent claims.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“The bug is fixed.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;How was it verified?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“The build works.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Which build command ran?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“There are no breaking changes.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What compatibility checks were performed?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“The migration is safe.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Was it tested against realistic data?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“This is secure.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What security checks were actually run?&lt;/p&gt;

&lt;p&gt;AI agents can sound extremely certain.&lt;/p&gt;

&lt;p&gt;Certainty is not evidence.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Better Completion Message
&lt;/h1&gt;

&lt;p&gt;Instead of this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Done.

Everything works and all tests pass.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I would rather see:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Implementation complete.

Verification performed:

npm test
Exit code: 0
128 tests passed
0 failed
3 skipped

npm run typecheck
Exit code: 0

npm run lint
Exit code: 0

Integration tests were NOT run.

Manual browser testing was NOT performed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is much more useful.&lt;/p&gt;

&lt;p&gt;Now I know exactly what was verified.&lt;/p&gt;

&lt;p&gt;And what was not.&lt;/p&gt;




&lt;h1&gt;
  
  
  Add “Not Verified” to Your Vocabulary
&lt;/h1&gt;

&lt;p&gt;One thing AI-assisted development needs more of is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Not verified.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is nothing wrong with an agent saying:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Implemented, but not tested.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unit tests pass, integration tests not run.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I could not verify this because the required service is unavailable.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those answers are better than pretending certainty.&lt;/p&gt;

&lt;p&gt;Good engineering is not about sounding confident.&lt;/p&gt;

&lt;p&gt;It is about knowing what evidence you actually have.&lt;/p&gt;




&lt;h1&gt;
  
  
  My Simple Rule
&lt;/h1&gt;

&lt;p&gt;For any important AI-generated change:&lt;/p&gt;

&lt;h2&gt;
  
  
  Never trust this:
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;“It works.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Prefer this:
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;“Here is how I verified it.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That one difference can prevent a lot of false confidence.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Practical Checklist
&lt;/h1&gt;

&lt;p&gt;Before accepting an agent's “tests pass” message, ask:&lt;/p&gt;

&lt;h3&gt;
  
  
  What command was executed?
&lt;/h3&gt;

&lt;p&gt;You should know the exact command.&lt;/p&gt;

&lt;h3&gt;
  
  
  Was it the full relevant test suite?
&lt;/h3&gt;

&lt;p&gt;Not just one subset.&lt;/p&gt;

&lt;h3&gt;
  
  
  What was the exit code?
&lt;/h3&gt;

&lt;p&gt;Success should be measurable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Were any tests skipped?
&lt;/h3&gt;

&lt;p&gt;Skipped tests matter.&lt;/p&gt;

&lt;h3&gt;
  
  
  Was verification run after the final code change?
&lt;/h3&gt;

&lt;p&gt;Not before the last edit.&lt;/p&gt;

&lt;h3&gt;
  
  
  Did CI confirm the result?
&lt;/h3&gt;

&lt;p&gt;Prefer independent verification.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I inspect the output?
&lt;/h3&gt;

&lt;p&gt;Evidence should be available.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;AI coding agents are becoming very good at producing code.&lt;/p&gt;

&lt;p&gt;But as they become more autonomous, developers need to become more careful about one thing:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;verification.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The dangerous workflow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI writes code
↓
AI says it works
↓
Human believes it
↓
Merge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The better workflow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI writes code
↓
AI provides evidence
↓
Independent checks run
↓
Human verifies
↓
Merge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Tests pass” is not evidence that tests passed.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is a claim.&lt;/p&gt;

&lt;p&gt;And in software engineering, important claims should come with receipts.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>discuss</category>
      <category>webdev</category>
    </item>
    <item>
      <title>If AI Writes the Code and AI Reviews the Code, What Exactly Is the Developer Verifying?</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Sat, 26 Sep 2026 03:59:59 +0000</pubDate>
      <link>https://dev.to/robertadam987_/if-ai-writes-the-code-and-ai-reviews-the-code-what-exactly-is-the-developer-verifying-b5h</link>
      <guid>https://dev.to/robertadam987_/if-ai-writes-the-code-and-ai-reviews-the-code-what-exactly-is-the-developer-verifying-b5h</guid>
      <description>&lt;p&gt;AI can now write the feature.&lt;/p&gt;

&lt;p&gt;Then AI can write the tests.&lt;/p&gt;

&lt;p&gt;Then AI can open the pull request.&lt;/p&gt;

&lt;p&gt;Then AI can review the pull request.&lt;/p&gt;

&lt;p&gt;Then AI can fix the review comments.&lt;/p&gt;

&lt;p&gt;And finally, a developer clicks:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Approve.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That workflow sounds incredibly efficient.&lt;/p&gt;

&lt;p&gt;It also creates a very uncomfortable question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If AI writes the code and AI reviews the code, what exactly is the human verifying?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is no longer theoretical.&lt;/p&gt;

&lt;p&gt;GitHub says Copilot code review now accounts for more than one in five code reviews on GitHub, and its review system can explore repository context, inspect large pull requests, review bot-authored PRs, and re-check its own findings after changes.&lt;/p&gt;

&lt;p&gt;That can be extremely useful.&lt;/p&gt;

&lt;p&gt;But only if we are clear about what the human reviewer is still responsible for.&lt;/p&gt;

&lt;p&gt;Because:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI reviewing AI-generated code is not the same thing as independent verification.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  The New Development Loop
&lt;/h1&gt;

&lt;p&gt;A growing number of workflows now look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Requirement
   ↓
AI generates implementation
   ↓
AI generates tests
   ↓
AI opens pull request
   ↓
AI reviews pull request
   ↓
AI fixes findings
   ↓
Human approves
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On paper, this looks great.&lt;/p&gt;

&lt;p&gt;Everything has been:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;implemented&lt;/li&gt;
&lt;li&gt;tested&lt;/li&gt;
&lt;li&gt;reviewed&lt;/li&gt;
&lt;li&gt;fixed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So what is left for the developer?&lt;/p&gt;

&lt;p&gt;A lot, actually.&lt;/p&gt;

&lt;p&gt;Because every step above can share the same misunderstanding.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Biggest Risk: Shared Assumptions
&lt;/h1&gt;

&lt;p&gt;Imagine the requirement is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A user should only receive a refund if the payment was successfully captured.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;AI misunderstands that as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A user can receive a refund if a payment record exists.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The AI then writes the implementation.&lt;/p&gt;

&lt;p&gt;Then it writes tests.&lt;/p&gt;

&lt;p&gt;The tests use the same interpretation.&lt;/p&gt;

&lt;p&gt;Then an AI reviewer inspects the code.&lt;/p&gt;

&lt;p&gt;It sees:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;clean structure&lt;/li&gt;
&lt;li&gt;tests passing&lt;/li&gt;
&lt;li&gt;correct types&lt;/li&gt;
&lt;li&gt;reasonable error handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything looks good.&lt;/p&gt;

&lt;p&gt;But the requirement is still wrong.&lt;/p&gt;

&lt;p&gt;The pipeline becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Wrong assumption
      ↓
Correct implementation of wrong assumption
      ↓
Correct tests for wrong assumption
      ↓
Correct review of wrong implementation
      ↓
Green CI
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is why passing tests and positive AI review are not enough.&lt;/p&gt;

&lt;p&gt;They can verify consistency.&lt;/p&gt;

&lt;p&gt;They cannot guarantee that the original understanding was correct.&lt;/p&gt;




&lt;h1&gt;
  
  
  AI Can Review Code Without Understanding Your Real Intent
&lt;/h1&gt;

&lt;p&gt;This is where human judgment still matters.&lt;/p&gt;

&lt;p&gt;An AI reviewer can often detect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;obvious bugs&lt;/li&gt;
&lt;li&gt;null handling&lt;/li&gt;
&lt;li&gt;unsafe patterns&lt;/li&gt;
&lt;li&gt;missing validation&lt;/li&gt;
&lt;li&gt;suspicious logic&lt;/li&gt;
&lt;li&gt;inconsistent naming&lt;/li&gt;
&lt;li&gt;duplicated code&lt;/li&gt;
&lt;li&gt;test gaps&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is valuable.&lt;/p&gt;

&lt;p&gt;GitHub has been expanding Copilot code review with deeper repository exploration, severity levels, and stronger agentic review capabilities.&lt;/p&gt;

&lt;p&gt;But your actual product intent may not live inside the codebase.&lt;/p&gt;

&lt;p&gt;It may live in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a conversation with a customer&lt;/li&gt;
&lt;li&gt;a product meeting&lt;/li&gt;
&lt;li&gt;a support ticket&lt;/li&gt;
&lt;li&gt;a legal requirement&lt;/li&gt;
&lt;li&gt;a business rule&lt;/li&gt;
&lt;li&gt;an undocumented edge case&lt;/li&gt;
&lt;li&gt;a decision made six months ago&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The model may never see that context.&lt;/p&gt;

&lt;p&gt;And even if you provide it, it can still misunderstand it.&lt;/p&gt;

&lt;p&gt;That means the human reviewer has to verify something deeper than syntax.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Developer Should Verify the Requirement First
&lt;/h1&gt;

&lt;p&gt;The first question should not be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Is this code good?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It should be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Is this solving the right problem?”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Before reviewing the implementation, ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What was the original requirement?&lt;/li&gt;
&lt;li&gt;What behavior should the user actually see?&lt;/li&gt;
&lt;li&gt;Which business rules matter?&lt;/li&gt;
&lt;li&gt;What should never happen?&lt;/li&gt;
&lt;li&gt;What assumptions did the AI make?&lt;/li&gt;
&lt;li&gt;Are those assumptions correct?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is requirement verification.&lt;/p&gt;

&lt;p&gt;And it may become one of the most important developer skills in AI-assisted development.&lt;/p&gt;




&lt;h1&gt;
  
  
  Tests Are Not Independent If AI Wrote Them From the Same Prompt
&lt;/h1&gt;

&lt;p&gt;This is another subtle problem.&lt;/p&gt;

&lt;p&gt;Suppose you give an AI:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Build a password-reset flow.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The AI writes the feature.&lt;/p&gt;

&lt;p&gt;Then you say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Write tests.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those tests are often based on the same mental model the AI used for the implementation.&lt;/p&gt;

&lt;p&gt;So if the model forgot an important requirement, the tests may forget it too.&lt;/p&gt;

&lt;p&gt;You get:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Implementation assumption
          ↓
Test assumption
          ↓
Both agree
          ↓
Tests pass
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is not independent verification.&lt;/p&gt;

&lt;p&gt;It is agreement.&lt;/p&gt;

&lt;p&gt;And agreement is not the same as correctness.&lt;/p&gt;




&lt;h1&gt;
  
  
  Ask Tests to Challenge the Implementation
&lt;/h1&gt;

&lt;p&gt;A better approach is to change the role of the test-generation step.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Write tests for this implementation.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;try:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Act as a skeptical QA engineer.

Do not assume the implementation is correct.

Based on the requirement, identify:
- failure cases
- abuse cases
- race conditions
- boundary cases
- invalid states
- unexpected user behavior

Then write tests that try to break the implementation.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That creates more separation between:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;creator&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;and&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;critic&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It is still AI reviewing AI.&lt;/p&gt;

&lt;p&gt;But at least you are changing the objective.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Human Reviewer Should Look for Decisions
&lt;/h1&gt;

&lt;p&gt;AI is very good at examining code.&lt;/p&gt;

&lt;p&gt;Humans should spend more time examining decisions.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Instead of asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is this function correct?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why does this function exist here?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does this API return the right type?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Should this endpoint be allowed to perform this operation at all?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is this query efficient?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Should this data be queried this way in the first place?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do the tests pass?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Are we testing the behavior the business actually needs?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is where developers add the most value.&lt;/p&gt;




&lt;h1&gt;
  
  
  Architecture Still Needs Human Attention
&lt;/h1&gt;

&lt;p&gt;AI reviewers can catch many local issues.&lt;/p&gt;

&lt;p&gt;But architecture is often about tradeoffs across the entire system.&lt;/p&gt;

&lt;p&gt;An AI-generated feature may introduce:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;another service&lt;/li&gt;
&lt;li&gt;another queue&lt;/li&gt;
&lt;li&gt;another database table&lt;/li&gt;
&lt;li&gt;another cache&lt;/li&gt;
&lt;li&gt;another dependency&lt;/li&gt;
&lt;li&gt;another abstraction&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each one can look reasonable individually.&lt;/p&gt;

&lt;p&gt;Together, they may make the system worse.&lt;/p&gt;

&lt;p&gt;That is why the human reviewer should ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Does this change make the system easier or harder to understand?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question is often more important than whether every line is technically correct.&lt;/p&gt;




&lt;h1&gt;
  
  
  Watch for “AI Agreement Loops”
&lt;/h1&gt;

&lt;p&gt;One pattern I think teams should actively avoid is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI writes code
↓
AI reviewer suggests change
↓
AI author accepts
↓
AI reviewer approves
↓
Human sees green check
↓
Merge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Everything in that loop may come from models with similar assumptions.&lt;/p&gt;

&lt;p&gt;The process looks highly reviewed.&lt;/p&gt;

&lt;p&gt;But the independence between steps may be weak.&lt;/p&gt;

&lt;p&gt;I think of this as an:&lt;/p&gt;

&lt;h1&gt;
  
  
  &lt;strong&gt;AI Agreement Loop&lt;/strong&gt;
&lt;/h1&gt;

&lt;p&gt;The system creates the appearance of verification because multiple stages agree.&lt;/p&gt;

&lt;p&gt;But those stages may not be truly independent.&lt;/p&gt;

&lt;p&gt;That is the key danger.&lt;/p&gt;




&lt;h1&gt;
  
  
  Green Checks Can Create False Confidence
&lt;/h1&gt;

&lt;p&gt;Developers are trained to trust signals like:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tests passed&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lint passed&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Type check passed&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI review passed&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security scan passed&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those signals are useful.&lt;/p&gt;

&lt;p&gt;But they can create a dangerous feeling:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Everything is green, so this must be safe.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Not necessarily.&lt;/p&gt;

&lt;p&gt;A green pipeline tells you:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The system passed the checks you chose to run.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It does not tell you:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;You chose the right checks.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That distinction matters more than ever.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Practical Human Review Checklist
&lt;/h1&gt;

&lt;p&gt;When AI writes the code and AI reviews the code, I think the developer should verify at least these 7 things.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Verify the Requirement
&lt;/h2&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does this actually solve what the user or business asked for?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does the code match the prompt?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those are different questions.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Verify the Assumptions
&lt;/h2&gt;

&lt;p&gt;Ask the AI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;List every assumption this implementation makes about:

- user behavior
- data
- permissions
- APIs
- infrastructure
- ordering
- timing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then review them manually.&lt;/p&gt;

&lt;p&gt;Hidden assumptions cause a huge number of production bugs.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Verify the Failure Modes
&lt;/h2&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What happens if the network fails?&lt;/li&gt;
&lt;li&gt;What happens if the DB write partially succeeds?&lt;/li&gt;
&lt;li&gt;What if the request is repeated?&lt;/li&gt;
&lt;li&gt;What if the external API times out?&lt;/li&gt;
&lt;li&gt;What if two requests happen simultaneously?&lt;/li&gt;
&lt;li&gt;What if the input is technically valid but unexpected?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Happy-path code is easy.&lt;/p&gt;

&lt;p&gt;Production failures live outside the happy path.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Verify the Blast Radius
&lt;/h2&gt;

&lt;p&gt;If this change fails, what breaks?&lt;/p&gt;

&lt;p&gt;One component?&lt;/p&gt;

&lt;p&gt;One customer?&lt;/p&gt;

&lt;p&gt;All customers?&lt;/p&gt;

&lt;p&gt;Payments?&lt;/p&gt;

&lt;p&gt;Authentication?&lt;/p&gt;

&lt;p&gt;Data integrity?&lt;/p&gt;

&lt;p&gt;A change touching a critical path deserves more human attention than a small UI change.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Verify the Architecture
&lt;/h2&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Did the AI add a new abstraction?&lt;/li&gt;
&lt;li&gt;Did it introduce another dependency?&lt;/li&gt;
&lt;li&gt;Did it duplicate existing functionality?&lt;/li&gt;
&lt;li&gt;Did it bypass an existing architectural boundary?&lt;/li&gt;
&lt;li&gt;Did it make the system harder to reason about?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Correct code can still create bad architecture.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Verify the Tests
&lt;/h2&gt;

&lt;p&gt;Do not only read whether tests pass.&lt;/p&gt;

&lt;p&gt;Read what they actually test.&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If the requirement were wrong, would these tests notice?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If the answer is no, you need better tests.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Verify That You Can Explain the Change
&lt;/h2&gt;

&lt;p&gt;This is my final rule.&lt;/p&gt;

&lt;p&gt;If someone asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Why was this implemented this way?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;you should have an answer better than:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“The agent generated it and Copilot approved it.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you cannot explain the change, you are not ready to own it.&lt;/p&gt;




&lt;h1&gt;
  
  
  Use Different Agents for Different Roles
&lt;/h1&gt;

&lt;p&gt;One useful pattern is to deliberately separate roles.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;h3&gt;
  
  
  Agent 1 — Implementer
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Implement the requirement using the existing architecture.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Agent 2 — Adversarial Reviewer
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Assume this implementation contains subtle bugs.

Try to find:
- incorrect assumptions
- security problems
- failure modes
- race conditions
- architectural problems
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Agent 3 — Test Designer
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Ignore the existing tests.

Design tests directly from the original requirement.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Human
&lt;/h3&gt;

&lt;p&gt;Verify:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Did all three understand the actual problem correctly?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Multiple agents do not replace human review.&lt;/p&gt;

&lt;p&gt;But role separation can reduce simple agreement loops.&lt;/p&gt;




&lt;h1&gt;
  
  
  Review the Requirement and the Diff Together
&lt;/h1&gt;

&lt;p&gt;One practical habit helps a lot.&lt;/p&gt;

&lt;p&gt;Do not review a pull request with only the code visible.&lt;/p&gt;

&lt;p&gt;Keep the original requirement next to it.&lt;/p&gt;

&lt;p&gt;For every significant change, compare:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Requirement
     ↕
Implementation
     ↕
Tests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These three should agree.&lt;/p&gt;

&lt;p&gt;If only:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Implementation ↔ Tests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;agree, you may still have a perfectly tested wrong feature.&lt;/p&gt;




&lt;h1&gt;
  
  
  Ask AI to Explain Before Asking It to Fix
&lt;/h1&gt;

&lt;p&gt;Suppose the AI reviewer finds a problem.&lt;/p&gt;

&lt;p&gt;Avoid immediately clicking:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Apply suggestion&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Explain:

1. Why this is a problem.
2. What real-world failure it could cause.
3. What assumptions the current implementation makes.
4. Why your proposed change fixes it.
5. What new risks your fix introduces.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now you are reviewing reasoning, not just accepting another generated diff.&lt;/p&gt;

&lt;p&gt;That helps maintain ownership.&lt;/p&gt;




&lt;h1&gt;
  
  
  AI Review Should Increase Human Leverage, Not Remove Human Judgment
&lt;/h1&gt;

&lt;p&gt;I am not arguing against AI code review.&lt;/p&gt;

&lt;p&gt;Quite the opposite.&lt;/p&gt;

&lt;p&gt;AI review can be extremely useful.&lt;/p&gt;

&lt;p&gt;It can catch issues humans miss.&lt;/p&gt;

&lt;p&gt;It can reduce repetitive review work.&lt;/p&gt;

&lt;p&gt;It can inspect large changes quickly.&lt;/p&gt;

&lt;p&gt;It can help teams focus attention where it matters.&lt;/p&gt;

&lt;p&gt;But the goal should be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI checks mechanics
+
Human checks meaning
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI writes
+
AI tests
+
AI reviews
+
Human clicks approve
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That second workflow may be fast.&lt;/p&gt;

&lt;p&gt;But eventually the developer becomes little more than the final button in an automated pipeline.&lt;/p&gt;

&lt;p&gt;And that is dangerous.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Human Role Is Moving Up a Level
&lt;/h1&gt;

&lt;p&gt;The developer's value is gradually shifting.&lt;/p&gt;

&lt;p&gt;Less time may go into:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;typing syntax&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;More time may go into:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;understanding requirements&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;designing systems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;challenging assumptions&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;verifying behavior&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;evaluating risk&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;debugging failures&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;making tradeoffs&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;taking ownership&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is not the disappearance of software engineering.&lt;/p&gt;

&lt;p&gt;It is software engineering becoming more about judgment.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Question I Now Ask Before Approving AI Code
&lt;/h1&gt;

&lt;p&gt;I used to ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Does this code look correct?”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now I think the better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“What do I know that the AI systems in this loop might not know?”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Maybe it is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;customer intent&lt;/li&gt;
&lt;li&gt;historical context&lt;/li&gt;
&lt;li&gt;production behavior&lt;/li&gt;
&lt;li&gt;an undocumented dependency&lt;/li&gt;
&lt;li&gt;an organizational constraint&lt;/li&gt;
&lt;li&gt;a previous incident&lt;/li&gt;
&lt;li&gt;a business rule&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That missing context may be the most important part of the review.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;We are quickly approaching a workflow where AI can:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;write the code&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;write the tests&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;review the code&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;fix the review&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;and maybe even:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;merge and deploy it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That does not make human verification less important.&lt;/p&gt;

&lt;p&gt;It changes what humans need to verify.&lt;/p&gt;

&lt;p&gt;The developer's job is increasingly not to check whether every semicolon is correct.&lt;/p&gt;

&lt;p&gt;It is to ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Did we build the right thing?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are the assumptions correct?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do we understand how it can fail?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can we explain why this design exists?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are we willing to own the consequences?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Because if AI writes the code and AI reviews the code, the most valuable thing the developer can provide is something neither system automatically guarantees:&lt;/p&gt;

&lt;h1&gt;
  
  
  &lt;strong&gt;Independent judgment.&lt;/strong&gt;
&lt;/h1&gt;

&lt;p&gt;AI can review the implementation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The developer still has to review reality.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>discuss</category>
    </item>
    <item>
      <title>AI Is Writing More of the Code — But Developers Are Becoming Responsible for More Than Ever</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Wed, 23 Sep 2026 03:44:46 +0000</pubDate>
      <link>https://dev.to/robertadam987_/ai-is-writing-more-of-the-code-but-developers-are-becoming-responsible-for-more-than-ever-55ni</link>
      <guid>https://dev.to/robertadam987_/ai-is-writing-more-of-the-code-but-developers-are-becoming-responsible-for-more-than-ever-55ni</guid>
      <description>&lt;p&gt;AI is writing more code.&lt;/p&gt;

&lt;p&gt;That part is obvious now.&lt;/p&gt;

&lt;p&gt;Developers can ask an agent to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;create a feature&lt;/li&gt;
&lt;li&gt;write tests&lt;/li&gt;
&lt;li&gt;fix a bug&lt;/li&gt;
&lt;li&gt;refactor a module&lt;/li&gt;
&lt;li&gt;update documentation&lt;/li&gt;
&lt;li&gt;generate SQL&lt;/li&gt;
&lt;li&gt;review a pull request&lt;/li&gt;
&lt;li&gt;even run commands and modify files&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And in many cases, it works surprisingly well.&lt;/p&gt;

&lt;p&gt;But something else is happening at the same time.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Developers are writing less of the code, while becoming responsible for more of it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the part I think we are underestimating.&lt;/p&gt;

&lt;p&gt;Because when production breaks at 2 AM, nobody is going to ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Which model generated this function?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;They are going to ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Who owns this system?”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And the answer is still us.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Old Workflow Was Easier to Reason About
&lt;/h2&gt;

&lt;p&gt;A traditional development workflow looked something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Understand requirement
        ↓
Design solution
        ↓
Write code
        ↓
Test it
        ↓
Review it
        ↓
Deploy it
        ↓
Maintain it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The developer touched almost every stage.&lt;/p&gt;

&lt;p&gt;That had problems.&lt;/p&gt;

&lt;p&gt;It was slower.&lt;/p&gt;

&lt;p&gt;It required more manual work.&lt;/p&gt;

&lt;p&gt;But there was one big advantage:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The person responsible for the code usually understood how it was created.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now the workflow is changing.&lt;/p&gt;




&lt;h1&gt;
  
  
  The New Workflow Looks More Like This
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Understand requirement
        ↓
Ask AI to plan
        ↓
AI writes implementation
        ↓
AI writes tests
        ↓
AI fixes errors
        ↓
AI reviews code
        ↓
Developer approves
        ↓
Production
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look carefully at where the human appears.&lt;/p&gt;

&lt;p&gt;Sometimes the developer is only deeply involved at the beginning and the end.&lt;/p&gt;

&lt;p&gt;That creates a strange new situation.&lt;/p&gt;

&lt;p&gt;The developer may be responsible for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;architecture they did not design&lt;/li&gt;
&lt;li&gt;code they did not write&lt;/li&gt;
&lt;li&gt;tests they did not write&lt;/li&gt;
&lt;li&gt;dependencies they did not choose&lt;/li&gt;
&lt;li&gt;edge cases they never considered&lt;/li&gt;
&lt;li&gt;behavior they only partially understand&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Yet the developer still owns the result.&lt;/p&gt;




&lt;h1&gt;
  
  
  Responsibility Did Not Disappear
&lt;/h1&gt;

&lt;p&gt;This is the mistake I think teams need to avoid.&lt;/p&gt;

&lt;p&gt;AI-generated code can make something feel less like your responsibility.&lt;/p&gt;

&lt;p&gt;It is easy to think:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“The agent generated it.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But production systems do not care who typed the code.&lt;/p&gt;

&lt;p&gt;If you merge it, approve it, or deploy it, it becomes part of the system you own.&lt;/p&gt;

&lt;p&gt;That means you are still responsible for questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is this secure?&lt;/li&gt;
&lt;li&gt;Does it actually meet the requirement?&lt;/li&gt;
&lt;li&gt;What happens if the API fails?&lt;/li&gt;
&lt;li&gt;Can this corrupt data?&lt;/li&gt;
&lt;li&gt;Does it create a race condition?&lt;/li&gt;
&lt;li&gt;Is the dependency trustworthy?&lt;/li&gt;
&lt;li&gt;Can another developer maintain it?&lt;/li&gt;
&lt;li&gt;What happens under load?&lt;/li&gt;
&lt;li&gt;What happens six months from now?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AI can help answer these questions.&lt;/p&gt;

&lt;p&gt;But it cannot remove your responsibility for them.&lt;/p&gt;




&lt;h1&gt;
  
  
  “The Tests Pass” Is Not Enough
&lt;/h1&gt;

&lt;p&gt;This becomes especially dangerous when AI writes both the code and the tests.&lt;/p&gt;

&lt;p&gt;Imagine this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Requirement
    ↓
AI misunderstands requirement
    ↓
AI writes implementation
    ↓
AI writes tests for its interpretation
    ↓
All tests pass
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Technically:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Green CI.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Practically:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wrong feature.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is one of the most important things developers need to remember in AI-assisted development.&lt;/p&gt;

&lt;p&gt;Tests can prove that code behaves according to a certain expectation.&lt;/p&gt;

&lt;p&gt;They cannot automatically prove that the expectation itself was correct.&lt;/p&gt;

&lt;p&gt;If the same model creates both the implementation and the test, they can share the same misunderstanding.&lt;/p&gt;

&lt;p&gt;That is why human review still matters at the requirement level.&lt;/p&gt;




&lt;h1&gt;
  
  
  Developers Are Becoming Supervisors of Software Creation
&lt;/h1&gt;

&lt;p&gt;The developer role is slowly changing from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Write
Test
Debug
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to something closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Define
Constrain
Delegate
Inspect
Verify
Approve
Own
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is not necessarily bad.&lt;/p&gt;

&lt;p&gt;In fact, it can be incredibly productive.&lt;/p&gt;

&lt;p&gt;But it requires a different skill set.&lt;/p&gt;

&lt;p&gt;The best developer may no longer be the person who can type code fastest.&lt;/p&gt;

&lt;p&gt;It may be the person who can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;define the problem clearly&lt;/li&gt;
&lt;li&gt;design good boundaries&lt;/li&gt;
&lt;li&gt;recognize bad architecture&lt;/li&gt;
&lt;li&gt;inspect a large diff quickly&lt;/li&gt;
&lt;li&gt;detect missing edge cases&lt;/li&gt;
&lt;li&gt;understand system behavior&lt;/li&gt;
&lt;li&gt;verify assumptions&lt;/li&gt;
&lt;li&gt;know when AI is confidently wrong&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are engineering skills, not typing skills.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Dangerous Part Is Losing System Understanding
&lt;/h1&gt;

&lt;p&gt;This is where I think the real risk begins.&lt;/p&gt;

&lt;p&gt;Suppose an AI agent adds a feature across 15 files.&lt;/p&gt;

&lt;p&gt;It modifies:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;database schema&lt;/li&gt;
&lt;li&gt;service logic&lt;/li&gt;
&lt;li&gt;API handlers&lt;/li&gt;
&lt;li&gt;validation&lt;/li&gt;
&lt;li&gt;background jobs&lt;/li&gt;
&lt;li&gt;caching&lt;/li&gt;
&lt;li&gt;tests&lt;/li&gt;
&lt;li&gt;frontend state&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything passes.&lt;/p&gt;

&lt;p&gt;You review the diff quickly.&lt;/p&gt;

&lt;p&gt;You merge.&lt;/p&gt;

&lt;p&gt;Then two months later, something breaks.&lt;/p&gt;

&lt;p&gt;Now you need to answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Why was this implemented this way?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But maybe nobody knows.&lt;/p&gt;

&lt;p&gt;The AI session is gone.&lt;/p&gt;

&lt;p&gt;The reasoning was never documented.&lt;/p&gt;

&lt;p&gt;The developer who approved it only reviewed the output.&lt;/p&gt;

&lt;p&gt;This is how teams can slowly build systems that &lt;strong&gt;work but are no longer deeply understood&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  Code Generation Can Create Understanding Debt
&lt;/h1&gt;

&lt;p&gt;We already talk about technical debt.&lt;/p&gt;

&lt;p&gt;But AI introduces something slightly different.&lt;/p&gt;

&lt;p&gt;I think of it as:&lt;/p&gt;

&lt;h1&gt;
  
  
  &lt;strong&gt;Understanding Debt&lt;/strong&gt;
&lt;/h1&gt;

&lt;p&gt;Technical debt is often:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“We know this code is messy, but we shipped it anyway.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Understanding debt is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“We own this code, but nobody fully understands why it works this way.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That can be even more dangerous.&lt;/p&gt;

&lt;p&gt;Because messy code is visible.&lt;/p&gt;

&lt;p&gt;Missing understanding is harder to detect.&lt;/p&gt;

&lt;p&gt;Everything may look fine until something unusual happens.&lt;/p&gt;




&lt;h1&gt;
  
  
  Small AI Tasks Are Safer Than Huge AI Tasks
&lt;/h1&gt;

&lt;p&gt;One simple way to reduce this problem is to stop giving agents enormous tasks.&lt;/p&gt;

&lt;p&gt;Compare these two prompts.&lt;/p&gt;

&lt;p&gt;Bad:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Build the full subscription system.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Better:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Analyze the current billing architecture.

Explain which files need to change.

Do not modify code yet.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Implement only the subscription data model.

Do not change unrelated files.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Add the billing service using the existing service pattern.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Add tests for the new behavior.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Smaller tasks create smaller diffs.&lt;/p&gt;

&lt;p&gt;Smaller diffs are easier to understand.&lt;/p&gt;

&lt;p&gt;And code that is easier to understand is easier to own.&lt;/p&gt;




&lt;h1&gt;
  
  
  Ask for the Plan Before the Code
&lt;/h1&gt;

&lt;p&gt;One of the most useful changes you can make is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Ask the AI to explain what it plans to do before it changes anything.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Before writing code:

1. Explain the current architecture.
2. List the files you plan to change.
3. Explain why each change is needed.
4. Identify possible risks.
5. Wait for approval before implementation.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives you a chance to catch bad direction early.&lt;/p&gt;

&lt;p&gt;Fixing a bad plan is cheap.&lt;/p&gt;

&lt;p&gt;Fixing 800 lines generated from a bad plan is not.&lt;/p&gt;




&lt;h1&gt;
  
  
  Review Decisions, Not Just Syntax
&lt;/h1&gt;

&lt;p&gt;Traditional code review often focuses on lines.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gi"&gt;+ const result = await processPayment()
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But with AI-generated code, the more important review may be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Why is payment processing happening here?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a different level of review.&lt;/p&gt;

&lt;p&gt;Instead of only asking:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is this line correct?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Should this responsibility live in this module?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does this function compile?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does this architecture make sense?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does the test pass?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is this testing the right behavior?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;AI is making syntax cheaper.&lt;/p&gt;

&lt;p&gt;That means developers need to spend more attention on decisions.&lt;/p&gt;




&lt;h1&gt;
  
  
  Never Merge Code You Cannot Explain
&lt;/h1&gt;

&lt;p&gt;This rule becomes much more important in the AI era.&lt;/p&gt;

&lt;p&gt;You do not need to memorize every line.&lt;/p&gt;

&lt;p&gt;But you should be able to explain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what changed&lt;/li&gt;
&lt;li&gt;why it changed&lt;/li&gt;
&lt;li&gt;what data moves through the system&lt;/li&gt;
&lt;li&gt;what can fail&lt;/li&gt;
&lt;li&gt;which assumptions exist&lt;/li&gt;
&lt;li&gt;which external services are involved&lt;/li&gt;
&lt;li&gt;how the feature can be rolled back&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If someone asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Why does this work this way?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and your answer is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“The AI generated it.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is not enough.&lt;/p&gt;

&lt;p&gt;Once it is merged, it is your system.&lt;/p&gt;




&lt;h1&gt;
  
  
  Let AI Generate Less Context Switching, Not More Complexity
&lt;/h1&gt;

&lt;p&gt;AI is extremely good at saving developers from repetitive work.&lt;/p&gt;

&lt;p&gt;That is where it can create huge value.&lt;/p&gt;

&lt;p&gt;Let it help with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;boilerplate&lt;/li&gt;
&lt;li&gt;repetitive tests&lt;/li&gt;
&lt;li&gt;simple migrations&lt;/li&gt;
&lt;li&gt;documentation&lt;/li&gt;
&lt;li&gt;refactoring suggestions&lt;/li&gt;
&lt;li&gt;test data&lt;/li&gt;
&lt;li&gt;type generation&lt;/li&gt;
&lt;li&gt;small bug fixes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But be careful when the AI starts generating:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;new architectural layers&lt;/li&gt;
&lt;li&gt;new dependencies&lt;/li&gt;
&lt;li&gt;large cross-cutting changes&lt;/li&gt;
&lt;li&gt;security-sensitive logic&lt;/li&gt;
&lt;li&gt;complicated concurrency&lt;/li&gt;
&lt;li&gt;data migrations&lt;/li&gt;
&lt;li&gt;payment logic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The higher the blast radius, the more human understanding should increase.&lt;/p&gt;

&lt;p&gt;Not decrease.&lt;/p&gt;




&lt;h1&gt;
  
  
  Documentation Matters More Now
&lt;/h1&gt;

&lt;p&gt;If AI made an important design choice, capture the reason.&lt;/p&gt;

&lt;p&gt;Not this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Process the payment&lt;/span&gt;
&lt;span class="nf"&gt;processPayment&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// We process payment before creating the final order because&lt;/span&gt;
&lt;span class="c1"&gt;// failed payments must not create confirmed inventory reservations.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That tells the next developer why the code exists.&lt;/p&gt;

&lt;p&gt;Architecture Decision Records can also help for larger choices.&lt;/p&gt;

&lt;p&gt;The AI session will disappear.&lt;/p&gt;

&lt;p&gt;The reasoning should not disappear with it.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Developer Still Needs to Know How to Debug
&lt;/h1&gt;

&lt;p&gt;There is another interesting effect here.&lt;/p&gt;

&lt;p&gt;The more code AI writes, the more valuable debugging becomes.&lt;/p&gt;

&lt;p&gt;Because when generated code fails, someone still needs to understand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;logs&lt;/li&gt;
&lt;li&gt;state&lt;/li&gt;
&lt;li&gt;timing&lt;/li&gt;
&lt;li&gt;network behavior&lt;/li&gt;
&lt;li&gt;databases&lt;/li&gt;
&lt;li&gt;dependencies&lt;/li&gt;
&lt;li&gt;concurrency&lt;/li&gt;
&lt;li&gt;deployment&lt;/li&gt;
&lt;li&gt;infrastructure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AI may help investigate.&lt;/p&gt;

&lt;p&gt;But a developer who does not understand the system will struggle to know whether the AI's explanation is correct.&lt;/p&gt;

&lt;p&gt;That is why I think debugging may become an even more important skill in the AI era.&lt;/p&gt;

&lt;p&gt;Writing code is becoming easier.&lt;/p&gt;

&lt;p&gt;Understanding why a system is broken is not.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Better AI Development Workflow
&lt;/h1&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Prompt
↓
Generate
↓
Tests pass
↓
Merge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Try this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Define requirement
        ↓
AI analyzes existing system
        ↓
AI proposes plan
        ↓
Developer reviews plan
        ↓
AI makes small change
        ↓
Developer understands diff
        ↓
Tests run
        ↓
Failure cases reviewed
        ↓
Architecture reviewed
        ↓
Merge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI is still doing a lot of work.&lt;/p&gt;

&lt;p&gt;But the developer remains connected to the reasoning.&lt;/p&gt;

&lt;p&gt;That is the important part.&lt;/p&gt;




&lt;h1&gt;
  
  
  Five Questions I Ask Before Approving AI-Generated Code
&lt;/h1&gt;

&lt;p&gt;Before merging, ask:&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Can I explain what changed?
&lt;/h2&gt;

&lt;p&gt;If not, keep reviewing.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Does this actually solve the requirement?
&lt;/h2&gt;

&lt;p&gt;Not just the test.&lt;/p&gt;

&lt;p&gt;The real requirement.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. What assumptions did the AI make?
&lt;/h2&gt;

&lt;p&gt;Look for hidden assumptions around data, APIs, users, permissions, and infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. What happens when something fails?
&lt;/h2&gt;

&lt;p&gt;Network timeout?&lt;/p&gt;

&lt;p&gt;Database error?&lt;/p&gt;

&lt;p&gt;Duplicate request?&lt;/p&gt;

&lt;p&gt;Partial write?&lt;/p&gt;

&lt;p&gt;Unexpected input?&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Will another developer understand this later?
&lt;/h2&gt;

&lt;p&gt;Because somebody will eventually maintain it.&lt;/p&gt;

&lt;p&gt;Maybe you.&lt;/p&gt;




&lt;h1&gt;
  
  
  AI Does Not Reduce Ownership
&lt;/h1&gt;

&lt;p&gt;This is probably the most important idea.&lt;/p&gt;

&lt;p&gt;AI can reduce:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;typing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;boilerplate&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;implementation time&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;repetitive work&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But it does not automatically reduce:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;responsibility&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;risk&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ownership&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;maintenance&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;production consequences&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In fact, AI may increase the amount of software a developer is responsible for.&lt;/p&gt;

&lt;p&gt;One engineer may soon oversee the amount of code that previously required several people to produce.&lt;/p&gt;

&lt;p&gt;That makes engineering judgment more important, not less.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Real Developer Skill Is Changing
&lt;/h1&gt;

&lt;p&gt;For years, software development was strongly associated with writing code.&lt;/p&gt;

&lt;p&gt;Now code itself is becoming easier to produce.&lt;/p&gt;

&lt;p&gt;So the valuable skills move upward.&lt;/p&gt;

&lt;p&gt;Understanding the problem.&lt;/p&gt;

&lt;p&gt;Choosing the architecture.&lt;/p&gt;

&lt;p&gt;Defining constraints.&lt;/p&gt;

&lt;p&gt;Recognizing risk.&lt;/p&gt;

&lt;p&gt;Reviewing decisions.&lt;/p&gt;

&lt;p&gt;Debugging failures.&lt;/p&gt;

&lt;p&gt;Protecting maintainability.&lt;/p&gt;

&lt;p&gt;Taking ownership.&lt;/p&gt;

&lt;p&gt;Those things are much harder to automate completely.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;AI may write more and more of our code.&lt;/p&gt;

&lt;p&gt;That is probably going to continue.&lt;/p&gt;

&lt;p&gt;But there is a dangerous assumption hiding behind that productivity:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If AI writes the code, AI owns the consequences.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It doesn't.&lt;/p&gt;

&lt;p&gt;You do.&lt;/p&gt;

&lt;p&gt;The future developer may write less code personally.&lt;/p&gt;

&lt;p&gt;But they may be responsible for &lt;strong&gt;more code, more systems, and more decisions than ever before&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So maybe the most important question in AI-assisted development is no longer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“Can the AI build this?”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“If I approve this, do I understand it well enough to own it?”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Because generating code is becoming cheap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Owning software is not.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>discuss</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How to Keep AI-Generated Code Maintainable After 6 Months</title>
      <dc:creator>Robert Adamson</dc:creator>
      <pubDate>Tue, 22 Sep 2026 04:48:18 +0000</pubDate>
      <link>https://dev.to/robertadam987_/how-to-keep-ai-generated-code-maintainable-after-6-months-3ha4</link>
      <guid>https://dev.to/robertadam987_/how-to-keep-ai-generated-code-maintainable-after-6-months-3ha4</guid>
      <description>&lt;p&gt;AI can help you write code incredibly fast.&lt;/p&gt;

&lt;p&gt;That part is no longer surprising.&lt;/p&gt;

&lt;p&gt;The harder question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Will you still understand that code six months from now?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is where many AI-assisted projects start to hurt.&lt;/p&gt;

&lt;p&gt;The code works today.&lt;/p&gt;

&lt;p&gt;Features ship quickly.&lt;/p&gt;

&lt;p&gt;Everything feels productive.&lt;/p&gt;

&lt;p&gt;Then a few months later:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;nobody knows why a function exists&lt;/li&gt;
&lt;li&gt;the same logic appears in three places&lt;/li&gt;
&lt;li&gt;naming is inconsistent&lt;/li&gt;
&lt;li&gt;one change breaks unrelated features&lt;/li&gt;
&lt;li&gt;tests are missing&lt;/li&gt;
&lt;li&gt;architecture becomes difficult to follow&lt;/li&gt;
&lt;li&gt;developers are afraid to refactor anything&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The problem is not that AI writes bad code every time.&lt;/p&gt;

&lt;p&gt;The problem is that AI is optimized to help you solve &lt;strong&gt;the current task&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Maintainability requires you to think about the codebase &lt;strong&gt;after hundreds of future tasks&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Here are 10 rules I use to keep AI-generated code maintainable.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Never Merge Code You Cannot Explain
&lt;/h2&gt;

&lt;p&gt;This is the most important rule.&lt;/p&gt;

&lt;p&gt;If AI generates 200 lines of code and you cannot explain what those 200 lines are doing, the job is not finished.&lt;/p&gt;

&lt;p&gt;You do not need to memorize every line.&lt;/p&gt;

&lt;p&gt;But you should understand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what the code does&lt;/li&gt;
&lt;li&gt;why it exists&lt;/li&gt;
&lt;li&gt;what inputs it expects&lt;/li&gt;
&lt;li&gt;what it changes&lt;/li&gt;
&lt;li&gt;what could fail&lt;/li&gt;
&lt;li&gt;what depends on it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A simple rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If you cannot explain the code to another developer, do not merge it yet.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ask the AI to explain the implementation if necessary.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Explain this implementation step by step.

Also tell me:

1. What assumptions does it make?
2. What could break?
3. Which parts are unnecessary?
4. Is there a simpler implementation?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AI should help you understand the code, not just generate more of it.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Ask for the Smallest Possible Change
&lt;/h2&gt;

&lt;p&gt;One of the easiest ways to destroy maintainability is allowing AI to modify too much at once.&lt;/p&gt;

&lt;p&gt;Imagine you ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Add user notifications.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An agent might:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;create new services&lt;/li&gt;
&lt;li&gt;modify database models&lt;/li&gt;
&lt;li&gt;change API routes&lt;/li&gt;
&lt;li&gt;add helper functions&lt;/li&gt;
&lt;li&gt;refactor unrelated files&lt;/li&gt;
&lt;li&gt;install another package&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The feature may work.&lt;/p&gt;

&lt;p&gt;But now reviewing it is much harder.&lt;/p&gt;

&lt;p&gt;Instead, break the work into smaller steps.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Step 1: Create the notification data model only.
Do not modify any other architecture.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Step 2: Add the notification service using the existing service pattern.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Step 3: Add the API endpoint.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Small changes are easier to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;understand&lt;/li&gt;
&lt;li&gt;review&lt;/li&gt;
&lt;li&gt;test&lt;/li&gt;
&lt;li&gt;revert&lt;/li&gt;
&lt;li&gt;debug&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AI can write code quickly.&lt;/p&gt;

&lt;p&gt;That does not mean you should let it change everything quickly.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Protect Your Existing Architecture
&lt;/h2&gt;

&lt;p&gt;AI does not always understand why your architecture looks the way it does.&lt;/p&gt;

&lt;p&gt;It may see:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;controllers/
services/
repositories/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and decide to introduce:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;managers/
handlers/
processors/
helpers/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now your project has two architectural styles.&lt;/p&gt;

&lt;p&gt;Six months later, nobody knows which one should be used.&lt;/p&gt;

&lt;p&gt;Before asking AI to implement something, give it architectural constraints.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Follow the existing architecture.

Controllers:
- validation and HTTP handling only

Services:
- business logic

Repositories:
- database operations

Do not introduce new architectural layers unless necessary.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This one instruction can prevent a lot of unnecessary complexity.&lt;/p&gt;

&lt;p&gt;Your AI should adapt to your codebase.&lt;/p&gt;

&lt;p&gt;Your codebase should not constantly adapt to your AI.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Make Naming Boring and Consistent
&lt;/h2&gt;

&lt;p&gt;AI often generates perfectly valid names that do not match the rest of the project.&lt;/p&gt;

&lt;p&gt;You may end up with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;getUser()
fetchUser()
retrieveUser()
loadUser()
findUser()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All performing similar operations.&lt;/p&gt;

&lt;p&gt;Individually, none of these names are wrong.&lt;/p&gt;

&lt;p&gt;Together, they create confusion.&lt;/p&gt;

&lt;p&gt;Maintainable projects usually have boring, predictable naming.&lt;/p&gt;

&lt;p&gt;If your project uses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;createUser
getUser
updateUser
deleteUser
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;keep using that pattern.&lt;/p&gt;

&lt;p&gt;Before generating code, tell the model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Follow the naming conventions already used in this repository.
Do not introduce new naming patterns.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Consistency is more valuable than creativity in production code.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Do Not Let AI Create Helpers for Everything
&lt;/h2&gt;

&lt;p&gt;AI loves abstraction.&lt;/p&gt;

&lt;p&gt;Sometimes too much.&lt;/p&gt;

&lt;p&gt;You ask for a small feature and suddenly you have:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UserProcessor
UserManager
UserHelper
UserFactory
UserTransformer
UserUtility
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;for something that could have been 20 lines inside an existing service.&lt;/p&gt;

&lt;p&gt;Abstraction is useful when it removes real duplication or complexity.&lt;/p&gt;

&lt;p&gt;It is harmful when it simply moves code into more files.&lt;/p&gt;

&lt;p&gt;Before accepting a new abstraction, ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What problem does this abstraction solve?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If the answer is only:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"It makes the code more reusable."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ask another question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where is it actually being reused?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If nowhere, you probably do not need it yet.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Keep Functions Small and Obvious
&lt;/h2&gt;

&lt;p&gt;AI can generate very large functions because it is trying to complete the entire task.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;createOrder&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// validate customer&lt;/span&gt;
  &lt;span class="c1"&gt;// check inventory&lt;/span&gt;
  &lt;span class="c1"&gt;// calculate discount&lt;/span&gt;
  &lt;span class="c1"&gt;// process payment&lt;/span&gt;
  &lt;span class="c1"&gt;// create order&lt;/span&gt;
  &lt;span class="c1"&gt;// update inventory&lt;/span&gt;
  &lt;span class="c1"&gt;// send email&lt;/span&gt;
  &lt;span class="c1"&gt;// create analytics event&lt;/span&gt;
  &lt;span class="c1"&gt;// notify admin&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It may work.&lt;/p&gt;

&lt;p&gt;But debugging it six months later will be painful.&lt;/p&gt;

&lt;p&gt;A better structure might be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;validateOrder&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;checkInventory&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;calculateTotal&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;processPayment&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;saveOrder&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;sendConfirmation&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each function has one clear responsibility.&lt;/p&gt;

&lt;p&gt;When asking AI to refactor, try:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Refactor this function into smaller functions.

Each function should have one clear responsibility.

Do not create unnecessary abstractions.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Simple code is easier for both humans and AI to work with later.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Make Tests Part of the Feature
&lt;/h2&gt;

&lt;p&gt;A common mistake is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI writes feature → developer checks UI → merge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That works until the next feature changes the same area.&lt;/p&gt;

&lt;p&gt;Then something silently breaks.&lt;/p&gt;

&lt;p&gt;Instead, make tests part of the original request.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Implement this feature and add tests for:

- expected behavior
- invalid input
- edge cases
- failure conditions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not treat tests as optional cleanup.&lt;/p&gt;

&lt;p&gt;They are documentation for future developers.&lt;/p&gt;

&lt;p&gt;Six months later, tests answer a very important question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What behavior was this code supposed to preserve?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That matters even more when much of the original implementation was generated by AI.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Ask AI to Remove Code, Not Only Add It
&lt;/h2&gt;

&lt;p&gt;AI development can create an interesting problem.&lt;/p&gt;

&lt;p&gt;Every prompt tends to add more code.&lt;/p&gt;

&lt;p&gt;Very few prompts remove anything.&lt;/p&gt;

&lt;p&gt;After months of development, you may accumulate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;unused helpers&lt;/li&gt;
&lt;li&gt;duplicate functions&lt;/li&gt;
&lt;li&gt;old feature flags&lt;/li&gt;
&lt;li&gt;abandoned interfaces&lt;/li&gt;
&lt;li&gt;unnecessary dependencies&lt;/li&gt;
&lt;li&gt;dead components&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Periodically ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Review this module for:

- duplicate logic
- dead code
- unnecessary abstractions
- unused dependencies
- functions that can be simplified

Do not change behavior.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is one of the best uses of AI.&lt;/p&gt;

&lt;p&gt;Use it not only as a code generator.&lt;/p&gt;

&lt;p&gt;Use it as a code cleaner.&lt;/p&gt;




&lt;h2&gt;
  
  
  9. Document Decisions, Not Obvious Code
&lt;/h2&gt;

&lt;p&gt;AI can generate comments everywhere:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// increment count&lt;/span&gt;
&lt;span class="nx"&gt;count&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That does not help anyone.&lt;/p&gt;

&lt;p&gt;Good documentation explains &lt;strong&gt;why&lt;/strong&gt; something exists.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// We intentionally retry only once here because the payment&lt;/span&gt;
&lt;span class="c1"&gt;// provider may create duplicate transactions on repeated requests.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That comment is useful.&lt;/p&gt;

&lt;p&gt;Six months later, a developer may otherwise "improve" the retry logic and create a serious bug.&lt;/p&gt;

&lt;p&gt;Document:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;unusual architectural decisions&lt;/li&gt;
&lt;li&gt;business rules&lt;/li&gt;
&lt;li&gt;important limitations&lt;/li&gt;
&lt;li&gt;third-party constraints&lt;/li&gt;
&lt;li&gt;performance tradeoffs&lt;/li&gt;
&lt;li&gt;security assumptions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not document things the code already makes obvious.&lt;/p&gt;




&lt;h2&gt;
  
  
  10. Review the Codebase Regularly
&lt;/h2&gt;

&lt;p&gt;Maintainability is not something you fix once.&lt;/p&gt;

&lt;p&gt;It slowly degrades.&lt;/p&gt;

&lt;p&gt;Especially when features are being generated quickly.&lt;/p&gt;

&lt;p&gt;Every few weeks, review the codebase and ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Where is complexity growing?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;files becoming too large&lt;/li&gt;
&lt;li&gt;repeated logic&lt;/li&gt;
&lt;li&gt;circular dependencies&lt;/li&gt;
&lt;li&gt;too many dependencies&lt;/li&gt;
&lt;li&gt;inconsistent patterns&lt;/li&gt;
&lt;li&gt;missing tests&lt;/li&gt;
&lt;li&gt;unclear responsibilities&lt;/li&gt;
&lt;li&gt;modules everyone is afraid to touch&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can even ask AI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Review this module as a senior engineer.

Do not rewrite it.

Identify maintainability problems that may become painful in 6–12 months.

Rank them by impact.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice the important part:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do not rewrite it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;First understand the problem.&lt;/p&gt;

&lt;p&gt;Then decide what should change.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Real Problem: AI Makes Bad Architecture Cheap
&lt;/h1&gt;

&lt;p&gt;Before AI, messy architecture took time to create.&lt;/p&gt;

&lt;p&gt;Now it can be created in minutes.&lt;/p&gt;

&lt;p&gt;That changes the economics of bad code.&lt;/p&gt;

&lt;p&gt;You can generate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;five new services&lt;/li&gt;
&lt;li&gt;ten interfaces&lt;/li&gt;
&lt;li&gt;three abstractions&lt;/li&gt;
&lt;li&gt;hundreds of lines of glue code&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;before you have really decided whether you need them.&lt;/p&gt;

&lt;p&gt;This means developers need to become more disciplined, not less.&lt;/p&gt;

&lt;p&gt;The bottleneck is no longer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Can we write this code?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Should this code exist in this form at all?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  A Better AI Coding Workflow
&lt;/h1&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Prompt
↓
Generate code
↓
Run it
↓
It works
↓
Merge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Try:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Define requirement
↓
Ask for implementation plan
↓
Review architecture
↓
Generate small change
↓
Understand the diff
↓
Run tests
↓
Review maintainability
↓
Merge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That may look slower.&lt;/p&gt;

&lt;p&gt;But it is much faster than debugging a codebase you no longer understand six months later.&lt;/p&gt;




&lt;h1&gt;
  
  
  My Final Checklist Before Merging AI-Generated Code
&lt;/h1&gt;

&lt;p&gt;Before merging, ask:&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I explain what this code does?
&lt;/h3&gt;

&lt;p&gt;If not, understand it first.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does it follow the existing architecture?
&lt;/h3&gt;

&lt;p&gt;If not, ask why.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is this the smallest reasonable change?
&lt;/h3&gt;

&lt;p&gt;Avoid unnecessary rewrites.&lt;/p&gt;

&lt;h3&gt;
  
  
  Did it introduce duplicate logic?
&lt;/h3&gt;

&lt;p&gt;Search before adding another helper.&lt;/p&gt;

&lt;h3&gt;
  
  
  Are the names consistent with the project?
&lt;/h3&gt;

&lt;p&gt;Consistency beats cleverness.&lt;/p&gt;

&lt;h3&gt;
  
  
  Are failure cases handled?
&lt;/h3&gt;

&lt;p&gt;Happy-path code is not enough.&lt;/p&gt;

&lt;h3&gt;
  
  
  Are tests included?
&lt;/h3&gt;

&lt;p&gt;Protect the behavior you just added.&lt;/p&gt;

&lt;h3&gt;
  
  
  Did it add a new dependency?
&lt;/h3&gt;

&lt;p&gt;Make sure you actually need it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is there a simpler solution?
&lt;/h3&gt;

&lt;p&gt;Ask this every time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Will another developer understand this in six months?
&lt;/h3&gt;

&lt;p&gt;That is the real test.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;AI makes writing software faster.&lt;/p&gt;

&lt;p&gt;But maintainable software has never been mainly about typing speed.&lt;/p&gt;

&lt;p&gt;It is about:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;clarity&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;consistency&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;architecture&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;testing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;good decisions&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;understanding tradeoffs&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;AI can generate thousands of lines for you.&lt;/p&gt;

&lt;p&gt;But those thousands of lines become &lt;strong&gt;your codebase&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And six months later, the AI that generated them may not remember why they were written.&lt;/p&gt;

&lt;p&gt;You and your team will still have to maintain them.&lt;/p&gt;

&lt;p&gt;So use AI to move faster.&lt;/p&gt;

&lt;p&gt;But keep one rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Never let your codebase grow faster than your understanding of it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>discuss</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
