<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ilya Golovenkov</title>
    <description>The latest articles on DEV Community by Ilya Golovenkov (@__f61d6e248).</description>
    <link>https://dev.to/__f61d6e248</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146369%2F91862d6c-106d-498c-9a1d-e41ca88d8615.jpg</url>
      <title>DEV Community: Ilya Golovenkov</title>
      <link>https://dev.to/__f61d6e248</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/__f61d6e248"/>
    <language>en</language>
    <item>
      <title>Why I stopped trusting “exit code 0” from AI coding agents</title>
      <dc:creator>Ilya Golovenkov</dc:creator>
      <pubDate>Mon, 28 Sep 2026 05:55:03 +0000</pubDate>
      <link>https://dev.to/__f61d6e248/why-i-stopped-trusting-exit-code-0-from-ai-coding-agents-1aaa</link>
      <guid>https://dev.to/__f61d6e248/why-i-stopped-trusting-exit-code-0-from-ai-coding-agents-1aaa</guid>
      <description>&lt;p&gt;AI coding agents are getting much better.&lt;/p&gt;

&lt;p&gt;They can inspect repositories, change multiple files, run commands, write tests, and sometimes complete work that would have taken a developer hours.&lt;/p&gt;

&lt;p&gt;But I kept running into the same problem:&lt;/p&gt;

&lt;p&gt;the agent would finish, return exit code 0, and say that everything was successful.&lt;/p&gt;

&lt;p&gt;Sometimes it really was.&lt;/p&gt;

&lt;p&gt;Sometimes it wasn't.&lt;/p&gt;

&lt;p&gt;That made me realize that I did not only need a better coding agent.&lt;/p&gt;

&lt;p&gt;I needed an independent way to verify what the agent had actually done.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;When an AI coding agent works on a project, several things can go wrong:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;required files may be missing before the run even starts;&lt;/li&gt;
&lt;li&gt;the agent may misunderstand the project structure;&lt;/li&gt;
&lt;li&gt;it may modify files it was not supposed to touch;&lt;/li&gt;
&lt;li&gt;tests may not actually run correctly;&lt;/li&gt;
&lt;li&gt;the agent may report success even when verification is incomplete;&lt;/li&gt;
&lt;li&gt;a successful process exit does not necessarily mean the project is actually in a good state.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The stronger coding agents become, the more important this problem becomes.&lt;/p&gt;

&lt;p&gt;If an agent changes 10 or 20 files, I do not want to manually inspect every line after every run.&lt;/p&gt;

&lt;p&gt;And if agents become even more autonomous, manual verification becomes even less practical.&lt;/p&gt;

&lt;p&gt;That is the problem I wanted to solve.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;I built a small open-source CLI called &lt;strong&gt;ELY Agent Input Preflight&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Try ELY Agent Input Preflight on GitHub:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/Golovenkov79/ely-agent-input-preflight" rel="noopener noreferrer"&gt;https://github.com/Golovenkov79/ely-agent-input-preflight&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The idea is simple:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Check whether the project is ready before the AI agent starts.&lt;/li&gt;
&lt;li&gt;Detect the project type.&lt;/li&gt;
&lt;li&gt;Prepare the right context and instructions.&lt;/li&gt;
&lt;li&gt;Run the coding agent in read-only mode by default.&lt;/li&gt;
&lt;li&gt;Monitor whether project files changed unexpectedly.&lt;/li&gt;
&lt;li&gt;Run independent verification after the agent finishes.&lt;/li&gt;
&lt;li&gt;Record the result locally.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The important part is this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;an agent exit code of 0 is not automatically treated as success.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;ELY can return:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;VERIFIED&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;FAILED&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;NEEDS_REVIEW&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the agent's own result is only one signal.&lt;/p&gt;

&lt;p&gt;It is not the final authority.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before the agent starts
&lt;/h2&gt;

&lt;p&gt;The first layer is Preflight.&lt;/p&gt;

&lt;p&gt;Before the coding agent runs, ELY checks whether the project has the inputs it is supposed to have.&lt;/p&gt;

&lt;p&gt;That can include things such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;required files;&lt;/li&gt;
&lt;li&gt;required directories;&lt;/li&gt;
&lt;li&gt;file patterns;&lt;/li&gt;
&lt;li&gt;environment variables;&lt;/li&gt;
&lt;li&gt;valid JSON files;&lt;/li&gt;
&lt;li&gt;non-empty files;&lt;/li&gt;
&lt;li&gt;optional files that should produce warnings instead of blocking the run.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If something critical is missing, ELY can stop the run before the coding agent starts.&lt;/p&gt;

&lt;p&gt;That sounds simple, but it avoids wasting an expensive agent run just to discover later that the required context was incomplete.&lt;/p&gt;

&lt;h2&gt;
  
  
  Project detection and skills
&lt;/h2&gt;

&lt;p&gt;ELY also detects the type of project it is looking at.&lt;/p&gt;

&lt;p&gt;The current release includes project detection for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Python;&lt;/li&gt;
&lt;li&gt;Android;&lt;/li&gt;
&lt;li&gt;.NET.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Based on that detection, ELY can select matching built-in guidance before preparing the handoff to the coding agent.&lt;/p&gt;

&lt;p&gt;The goal is not to replace the agent.&lt;/p&gt;

&lt;p&gt;The goal is to make sure the agent starts with better context and clearer rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  Read-only by default
&lt;/h2&gt;

&lt;p&gt;One design decision was especially important to me:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the coding agent should not get write access by default.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;With the current Codex integration, ELY starts the agent in an explicit read-only sandbox unless write access is deliberately enabled.&lt;/p&gt;

&lt;p&gt;If I only want an agent to inspect a project, I do not want it silently modifying files.&lt;/p&gt;

&lt;p&gt;Write access has to be explicitly requested.&lt;/p&gt;

&lt;p&gt;That gives me a much clearer separation between:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;inspection;&lt;/li&gt;
&lt;li&gt;modification.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  File integrity monitoring
&lt;/h2&gt;

&lt;p&gt;ELY also takes a snapshot of project files before an agent run.&lt;/p&gt;

&lt;p&gt;After the agent exits, it compares the workspace again.&lt;/p&gt;

&lt;p&gt;This allows ELY to detect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;created files;&lt;/li&gt;
&lt;li&gt;modified files;&lt;/li&gt;
&lt;li&gt;deleted files.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That matters especially for read-only runs.&lt;/p&gt;

&lt;p&gt;If a supposedly read-only agent run changes protected project files, ELY can treat the result as a failure.&lt;/p&gt;

&lt;p&gt;This is independent of what the agent claims it did.&lt;/p&gt;

&lt;h2&gt;
  
  
  A real example
&lt;/h2&gt;

&lt;p&gt;During development I had a Codex run where the agent completed with exit code 0.&lt;/p&gt;

&lt;p&gt;Inside its sandbox, however, many tests could not run correctly because temporary writable storage was unavailable.&lt;/p&gt;

&lt;p&gt;The agent still completed its inspection and reported the limitation.&lt;/p&gt;

&lt;p&gt;ELY did not simply trust the agent result.&lt;/p&gt;

&lt;p&gt;After Codex exited, ELY ran its own Postflight verification outside the agent sandbox and independently checked the project.&lt;/p&gt;

&lt;p&gt;The Postflight verifier was able to run the project's tests normally and verify the state independently.&lt;/p&gt;

&lt;p&gt;That was the moment when the idea became useful to me.&lt;/p&gt;

&lt;p&gt;The interesting part was not that the coding agent was "bad".&lt;/p&gt;

&lt;p&gt;It was that the agent and the verifier were operating under different conditions.&lt;/p&gt;

&lt;p&gt;A successful agent run and a successfully verified project are not always the same thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Independent Postflight verification
&lt;/h2&gt;

&lt;p&gt;After the coding agent exits, ELY performs its own verification step.&lt;/p&gt;

&lt;p&gt;For Python projects, the current built-in verifier can run the project's unittest suite when a &lt;code&gt;tests&lt;/code&gt; directory is available.&lt;/p&gt;

&lt;p&gt;The final result can then be classified as:&lt;/p&gt;

&lt;h3&gt;
  
  
  VERIFIED
&lt;/h3&gt;

&lt;p&gt;The independent verification passed.&lt;/p&gt;

&lt;h3&gt;
  
  
  FAILED
&lt;/h3&gt;

&lt;p&gt;The agent failed, the project changed when it should not have, the Preflight checks became blocked, or the independent verifier failed.&lt;/p&gt;

&lt;h3&gt;
  
  
  NEEDS_REVIEW
&lt;/h3&gt;

&lt;p&gt;ELY does not currently have enough automated verification for that project type.&lt;/p&gt;

&lt;p&gt;I prefer this to pretending that every run can be automatically proven correct.&lt;/p&gt;

&lt;p&gt;Sometimes the right answer really is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;this needs a human review.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Current flow
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
text
Preflight
  -&amp;gt; Project detection
  -&amp;gt; Skill selection
  -&amp;gt; Agent handoff
  -&amp;gt; Codex
  -&amp;gt; File integrity check
  -&amp;gt; Independent Postflight verification
  -&amp;gt; Run history
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>testing</category>
      <category>python</category>
    </item>
  </channel>
</rss>
