A few days ago I wrote about why I stopped treating exit code 0 from an AI coding agent as proof that the requested work was actually completed correctly.
That idea has now turned into the next release of ELY Agent Input Preflight: v0.7.0.
The biggest change is a new layer called Generated-Code Admission.
The goal is simple:
before generated source code is allowed to contribute to a successful final result, ELY should independently check whether that source can be admitted at all.
GitHub:
https://github.com/Golovenkov79/ely-agent-input-preflight
Release v0.7.0:
https://github.com/Golovenkov79/ely-agent-input-preflight/releases/tag/v0.7.0
Project page:
https://elyzorix.com/tools/agent-input-preflight/
The verification boundary is now more explicit
One of the most useful pieces of feedback on my previous post was that there are several different questions hidden behind a single word like "success".
Did the agent process complete?
Did the requested operations actually execute?
Does the resulting repository satisfy the acceptance criteria?
Can the generated source itself be independently admitted?
I now keep those questions separate:
PROCESS
↓
EXECUTION
↓
ACCEPTANCE
↓
FINAL
The important part is that these states are not silently collapsed into one green check.
An agent can complete successfully while the repository still needs review.
A test command can run while the generated source itself is invalid or out of scope.
And sometimes the honest answer is simply:
NOT_PROVEN
which then becomes:
NEEDS_REVIEW
instead of a false VERIFIED.
Generated-Code Admission
v0.7.0 adds an independent admission layer between AI-generated source changes and the final Postflight decision.
For Python source, the Community release currently performs bounded, deterministic checks including:
- syntax validation;
- structural parsing;
- source-scope checks;
- project-local policy checks;
- baseline dangerous-construct detection.
The baseline currently detects or blocks supported forms of:
eval(...)
exec(...)
os.system(...)
subprocess.run(..., shell=True)
including supported import aliases.
This is intentionally not presented as a complete security scanner.
It is a bounded admission layer with deterministic rules.
If ELY cannot prove a result, it does not silently convert uncertainty into safety.
For example, a dynamic subprocess flag such as:
subprocess.run(cmd, shell=flag)
cannot automatically be treated as safe just because flag is not literally True.
The admission result stays NOT_PROVEN.
Project-local scope and policy
A project can also define local admission rules.
For example:
{
"admission": {
"allowed_paths": ["src/*.py"],
"denied_paths": ["secrets/*"],
"forbidden_calls": ["open"]
}
}
This lets the project say not only "is this valid Python?", but also:
- was this file supposed to be changed here?
- is this path allowed?
- is this call allowed by this project's policy?
An invalid, dangerous, out-of-scope, or policy-violating source does not get promoted to a successful final result.
Why I did not put the safety layer behind Pro
I also made a product decision while building this release.
The Community version keeps the full core verification and safety functionality.
I do not want a model where:
"we found a dangerous generated-code pattern, but you need the paid plan to verify it properly."
That would defeat the purpose of the project.
The future Pro layer is planned around time and automation, not better safety.
Examples include:
- one-command end-to-end workflows;
- reusable workflow profiles;
- batch and multi-project orchestration;
- CI / PR automation;
- richer reports;
- history analytics;
- unattended or scheduled workflows;
- team-oriented workflow management later.
The same verification core should remain available in Community.
A short version of the product rule is:
Pro sells time, not safety.
Local-first is still the default
The Community core remains local-first.
It does not require:
- an ELY account;
- a remote dashboard;
- telemetry;
- a cloud database.
Run history and verification evidence remain local.
The project is still released under Apache-2.0.
What v0.7.0 was tested against
Before publishing the release, I ran the full regression suite plus adversarial admission cases.
The final release candidate passed:
- 169 automated tests;
- self-preflight;
- release archive audit;
- SHA-256 verification;
- fresh Windows install smoke;
- successful admission smoke;
- adversarial
NOT_PROVENsmoke.
The Windows source package SHA-256 is:
698996db4534579dc99b5c4625e3c56ed246b7e188117919e65ddc75bfd351cf
The current idea
The project started with a simple rule:
an AI agent saying "success" is evidence, not authority.
v0.7.0 pushes that idea one step further.
Now the generated source itself has to cross an independent admission boundary before it can help produce a verified final result.
There is still a lot to build.
But I would rather have a system that says:
NEEDS_REVIEW
when evidence is incomplete than one that produces a confident green check from assumptions.
If you work with coding agents and have run into similar verification problems, I would be interested in hearing where your current workflow still relies on trust rather than independent evidence.
GitHub:
https://github.com/Golovenkov79/ely-agent-input-preflight
Release:
https://github.com/Golovenkov79/ely-agent-input-preflight/releases/tag/v0.7.0
Top comments (0)