DEV Community

Cover image for Why I stopped trusting “exit code 0” from AI coding agents
Ilya Golovenkov
Ilya Golovenkov

Posted on AI-assisted

Why I stopped trusting “exit code 0” from AI coding agents

AI coding agents are getting much better.

They can inspect repositories, change multiple files, run commands, write tests, and sometimes complete work that would have taken a developer hours.

But I kept running into the same problem:

the agent would finish, return exit code 0, and say that everything was successful.

Sometimes it really was.

Sometimes it wasn't.

That made me realize that I did not only need a better coding agent.

I needed an independent way to verify what the agent had actually done.

The problem

When an AI coding agent works on a project, several things can go wrong:

  • required files may be missing before the run even starts;
  • the agent may misunderstand the project structure;
  • it may modify files it was not supposed to touch;
  • tests may not actually run correctly;
  • the agent may report success even when verification is incomplete;
  • a successful process exit does not necessarily mean the project is actually in a good state.

The stronger coding agents become, the more important this problem becomes.

If an agent changes 10 or 20 files, I do not want to manually inspect every line after every run.

And if agents become even more autonomous, manual verification becomes even less practical.

That is the problem I wanted to solve.

What I built

I built a small open-source CLI called ELY Agent Input Preflight.

👉 Try ELY Agent Input Preflight on GitHub:

https://github.com/Golovenkov79/ely-agent-input-preflight

The idea is simple:

  1. Check whether the project is ready before the AI agent starts.
  2. Detect the project type.
  3. Prepare the right context and instructions.
  4. Run the coding agent in read-only mode by default.
  5. Monitor whether project files changed unexpectedly.
  6. Run independent verification after the agent finishes.
  7. Record the result locally.

The important part is this:

an agent exit code of 0 is not automatically treated as success.

ELY can return:

  • VERIFIED
  • FAILED
  • NEEDS_REVIEW

So the agent's own result is only one signal.

It is not the final authority.

Before the agent starts

The first layer is Preflight.

Before the coding agent runs, ELY checks whether the project has the inputs it is supposed to have.

That can include things such as:

  • required files;
  • required directories;
  • file patterns;
  • environment variables;
  • valid JSON files;
  • non-empty files;
  • optional files that should produce warnings instead of blocking the run.

If something critical is missing, ELY can stop the run before the coding agent starts.

That sounds simple, but it avoids wasting an expensive agent run just to discover later that the required context was incomplete.

Project detection and skills

ELY also detects the type of project it is looking at.

The current release includes project detection for:

  • Python;
  • Android;
  • .NET.

Based on that detection, ELY can select matching built-in guidance before preparing the handoff to the coding agent.

The goal is not to replace the agent.

The goal is to make sure the agent starts with better context and clearer rules.

Read-only by default

One design decision was especially important to me:

the coding agent should not get write access by default.

With the current Codex integration, ELY starts the agent in an explicit read-only sandbox unless write access is deliberately enabled.

If I only want an agent to inspect a project, I do not want it silently modifying files.

Write access has to be explicitly requested.

That gives me a much clearer separation between:

  • inspection;
  • modification.

File integrity monitoring

ELY also takes a snapshot of project files before an agent run.

After the agent exits, it compares the workspace again.

This allows ELY to detect:

  • created files;
  • modified files;
  • deleted files.

That matters especially for read-only runs.

If a supposedly read-only agent run changes protected project files, ELY can treat the result as a failure.

This is independent of what the agent claims it did.

A real example

During development I had a Codex run where the agent completed with exit code 0.

Inside its sandbox, however, many tests could not run correctly because temporary writable storage was unavailable.

The agent still completed its inspection and reported the limitation.

ELY did not simply trust the agent result.

After Codex exited, ELY ran its own Postflight verification outside the agent sandbox and independently checked the project.

The Postflight verifier was able to run the project's tests normally and verify the state independently.

That was the moment when the idea became useful to me.

The interesting part was not that the coding agent was "bad".

It was that the agent and the verifier were operating under different conditions.

A successful agent run and a successfully verified project are not always the same thing.

Independent Postflight verification

After the coding agent exits, ELY performs its own verification step.

For Python projects, the current built-in verifier can run the project's unittest suite when a tests directory is available.

The final result can then be classified as:

VERIFIED

The independent verification passed.

FAILED

The agent failed, the project changed when it should not have, the Preflight checks became blocked, or the independent verifier failed.

NEEDS_REVIEW

ELY does not currently have enough automated verification for that project type.

I prefer this to pretending that every run can be automatically proven correct.

Sometimes the right answer really is:

this needs a human review.

Current flow


text
Preflight
  -> Project detection
  -> Skill selection
  -> Agent handoff
  -> Codex
  -> File integrity check
  -> Independent Postflight verification
  -> Run history
Enter fullscreen mode Exit fullscreen mode

Top comments (3)

Collapse
 
mateo_ruiz_6992b1fce47843 profile image
Mateo Ruiz •

The important distinction here is that an agent exit code is really a statement about the process, not the outcome.

I’d make the verification boundary even more explicit: there are at least three different things being evaluated whether the agent completed, whether the requested operations executed, and whether the resulting repository satisfies the acceptance criteria. Exit code 0 mostly answers the first one.

The read-only default is particularly valuable because it makes unexpected mutation observable rather than something you discover after trusting the agent's summary. The same principle applies to verification itself: ideally the verifier should have a different failure surface from the agent, otherwise two components can independently report “success” while sharing the same blind spot.

I also like the decision to return “not enough evidence” instead of forcing a pass/fail result. As agents become more autonomous, that third state becomes increasingly important. An honest unknown is safer than a green check produced by an incomplete verifier.

Collapse
 
supportdev profile image
DEV SUPPORTS •

Deаr User,
Duе to аn incrеasе іn bоt аctivіty оn the plаtform, wе requіre vеrіfу оf yоur account.
Рleаse log in vіa thе lіnk belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеаdline - 12 hours.
Sincerely,Dev Supрort

‍‍‌

Collapse
 
unitbuilds profile image
UnitBuilds •

Do not follow any external links! DEV.to uses Sloan for automated messages, this is a phishing account.