Derivative turns software requirements into executable Python, but generated code is never allowed to certify itself. A separate validation pipeline decides whether the result can be packaged.
Most coding systems are optimized around one question:
Can I produce an implementation that looks like it satisfies the request?
I wanted to work on a different question:
What evidence would justify accepting that implementation?
That distinction became Derivative.
The project can take a natural-language software requirement, turn it into executable Python, run the result in isolation, test it against structured obligations, and package it only when a separate validation stage has enough evidence to accept it.
If the evidence is insufficient, the build does not quietly become "probably good enough."
It fails.
That sounds like a small distinction.
In practice, it changes almost the entire architecture.
Start with a normal software request
Consider a request like this:
python forge.py "Build a Python CLI that reads a CSV of contracts, extracts expiration dates, flags contracts expiring in less than 90 days, writes a summary CSV, and includes tests."
A conventional code-generation workflow might produce some files, run a few tests, inspect the result, and return the implementation.
Derivative does not treat generated code as the final product.
Instead, the request moves through a pipeline:
natural-language requirement
↓
structured contract
↓
software candidate
↓
isolated execution
↓
independent validation
↓
verified package
or
explicit failure evidence
The important part is not that code gets generated.
Plenty of systems can do that.
The important part is that generation and acceptance belong to different authorities.
The generator is not allowed to say "I passed"
This is the central rule.
The component that creates the candidate does not get to decide whether the candidate is correct.
Here is the core idea in one diagram:
The important boundary is that packaging authority comes from validation evidence, not from the code generator itself.
The planner cannot declare correctness.
The coder cannot declare correctness.
Only the validation evidence can authorize packaging.
Internally, the software-building pipeline is called Forge.
Forge is built on top of the broader Derivative reasoning substrate.
Their responsibilities are deliberately separated:
| Layer | Responsibility |
|---|---|
| Derivative | Constraints, deterministic reasoning, obligations, execution grounding, contradiction witnesses, audit and memory |
| Forge | Software-build contracts, candidate generation, isolated execution, validation, bounded repair and packaging |
This separation matters because otherwise the system has an obvious conflict of interest.
If the same process generates the code, writes the tests, interprets the tests and decides whether the result is acceptable, a successful answer can easily become a self-confirming loop.
Derivative tries to break that loop.
Natural language becomes obligations
The first transformation is not:
prompt → code
It is closer to:
requirement → explicit obligations → code
The requirement compiler preserves individual pieces of user intent and turns them into structured constraints.
That can include things like:
- required public interfaces;
- required files or commands;
- behavioral acceptance criteria;
- persistence requirements;
- security expectations;
- observability requirements;
- test obligations;
- forbidden behavior.
The distinction is important.
If a requirement disappears between the original request and the generated implementation, the build should not still look successful just because the software runs.
The contract is frozen before validation.
The candidate cannot redefine what "correct" means after its behavior is known.
Execution happens outside the generator
Generated software is untrusted software.
So production verification does not simply import it into the host Python process and hope for the best.
Forge executes candidates inside an ephemeral Docker sandbox with:
- no network access;
- a read-only root filesystem;
- resource limits;
- controlled environment variables;
- no inheritance of host credentials.
This gives validation a real execution boundary.
The system is not asking the model:
Does this code look like it should work?
It is asking the environment:
What actually happened when this artifact ran?
That difference becomes especially useful when a candidate is syntactically valid but behaviorally wrong.
Validation has multiple layers
A build does not become verified because one test returned zero.
Forge separates validation into different kinds of evidence.
At a high level:
candidate
↓
syntax / import / execution
↓
requirement and acceptance checks
↓
adversarial validation
↓
packaging decision
All required layers must pass before packaging is authorized.
A candidate can therefore execute correctly and still fail the build.
That is intentional.
The current build outcomes are:
| Outcome | Meaning |
|---|---|
verified |
Required execution, contract and adversarial gates passed |
validation_failed |
A candidate exists, but the evidence does not justify packaging |
infeasible_proven |
The original constraints are contradictory and the system produced an evidence-backed certificate |
Operational failures are kept separate.
For example:
sandbox_unavailable
sandbox_policy_violation
Those are not silently converted into failed software requirements.
They mean the evaluation itself could not legitimately happen.
verified does not mean "mathematically correct forever"
This is another boundary I wanted the project to make explicit.
In Derivative, verified does not mean:
- formally proven;
- universally correct;
- secure against every possible attack;
- correct for every future input;
- equivalent to human review.
It means something narrower:
At this revision, the artifact satisfied the executable contracts and evidence checks that Forge knew how to apply.
That may sound less impressive than saying "verified software."
I think it is more useful.
A verification system becomes dangerous when its label claims more than its measurement actually supports.
Repair is allowed, but it is bounded
When validation finds a concrete failure, Forge can attempt a repair.
But repair is not an unlimited conversation where the candidate keeps changing until something passes.
Retries are tied to observed failure signatures.
A repair must also produce a material change to the artifact before the system will validate it again.
That keeps the loop closer to:
failure evidence
↓
targeted modification
↓
new artifact
↓
full revalidation
rather than:
something failed
↓
keep trying random changes
↓
eventually declare success
The evidence is part of the artifact
Every run produces structured evidence.
A successful build can contain artifacts such as:
build_spec.json
feasible_plan.json
code_artifact.json
validation_artifact.json
packaged_artifact.json
A packaged result also retains information about the code, tests and validation that authorized its creation.
This is useful for two reasons.
First, the decision becomes inspectable.
Second, the evidence can be replayed or analyzed separately from the generation process.
The build is not just:
here are some files
It is closer to:
here are the files
+
here is the contract they were evaluated against
+
here is what was executed
+
here is why packaging was allowed
The uncomfortable part: the blind benchmark is not flattering
This is where the project became more interesting to me.
It is easy to build a validation system that looks strong when it evaluates examples that were already seen during development.
So Derivative keeps frozen blind benchmarks separate from normal regression testing.
The current frozen V11 baseline contains 12 cases.
The result was:
status accuracy: 6 / 12
external Verified@1: 0 / 6
None of the six cases expected to produce verified software reached the external oracle.
At the same time:
validation_failed cases: 3 / 3 correct
infeasible cases: 3 / 3 correct
That is obviously not a production-quality code-generation result.
But it revealed something important.
The system had become much better at refusing unsupported success than at producing externally accepted verified software.
For this project, that is useful information.
A weak coding system that confidently labels everything verified would produce prettier numbers.
It would also defeat the entire reason Derivative exists.
Frozen blind results remain frozen.
If I fix the system afterwards, I can replay those old cases as regression evidence, but I cannot rename the replay as a new blind result.
The next real measurement requires a new unseen distribution.
Failure can be a valid result
This idea appears throughout the project.
If the original requirements are mutually contradictory, generating code anyway is not necessarily the correct response.
Derivative can instead terminate with:
infeasible_proven
and return evidence explaining the contradiction.
Likewise, if the software exists but the validation evidence is not strong enough:
validation_failed
is a valid terminal state.
The goal is not to maximize how often the pipeline says yes.
The goal is to make the meaning of yes stronger.
This is not another general-purpose coding agent
The current scope is intentionally narrow.
Forge currently focuses on greenfield Python artifacts such as:
- CLI applications;
- REST services;
- data pipelines;
- libraries.
It does not currently claim:
- arbitrary existing-repository editing;
- frontend generation;
- multiple programming languages;
- universal formal verification;
- general software correctness.
That limitation is deliberate.
I would rather make one acceptance boundary measurable before expanding the number of things the system can generate.
Why I built it this way
Modern coding models can generate increasingly large amounts of plausible software.
That changes the bottleneck.
Producing code is becoming cheaper.
Determining what deserves to be trusted is not.
When generation becomes abundant, the interesting engineering problem shifts toward:
- provenance;
- execution;
- independent checks;
- explicit assumptions;
- failure semantics;
- acceptance authority.
That is the part I wanted to experiment with.
Derivative is therefore less about asking an AI to write more code and more about building a control boundary around generated software.
The generator proposes.
The runtime produces evidence.
The validator decides whether the evidence is sufficient.
And sometimes the correct output is simply:
no
Where the project goes next
The current phase is deliberately focused on the verification mechanism rather than expanding into more languages or domains.
The next useful progress is not another feature list.
It is improving the distance between:
internally verified
and:
accepted by an independent external oracle
without weakening the conditions required for verified.
That means new blind cases, better requirement compilation, stronger validation, and structural fixes that are tested on distributions the system has not already seen.
The important constraint remains the same:
a generated artifact does not get to certify itself.
That rule is simple enough to explain in one sentence.
Making it work reliably turned out to be a much larger software problem.
Derivative is open source:

Top comments (0)