DEV Community

Cover image for A Coding System That Refuses to Trust Its Own Output
DaC
DaC

Posted on

A Coding System That Refuses to Trust Its Own Output

Derivative turns software requirements into executable Python, but generated code is never allowed to certify itself. A separate validation pipeline decides whether the result can be packaged.

Most coding systems are optimized around one question:

Can I produce an implementation that looks like it satisfies the request?

I wanted to work on a different question:

What evidence would justify accepting that implementation?

That distinction became Derivative.

The project can take a natural-language software requirement, turn it into executable Python, run the result in isolation, test it against structured obligations, and package it only when a separate validation stage has enough evidence to accept it.

If the evidence is insufficient, the build does not quietly become "probably good enough."

It fails.

That sounds like a small distinction.

In practice, it changes almost the entire architecture.


Start with a normal software request

Consider a request like this:

python forge.py "Build a Python CLI that reads a CSV of contracts, extracts expiration dates, flags contracts expiring in less than 90 days, writes a summary CSV, and includes tests."
Enter fullscreen mode Exit fullscreen mode

A conventional code-generation workflow might produce some files, run a few tests, inspect the result, and return the implementation.

Derivative does not treat generated code as the final product.

Instead, the request moves through a pipeline:

natural-language requirement
        ↓
structured contract
        ↓
software candidate
        ↓
isolated execution
        ↓
independent validation
        ↓
verified package
        or
explicit failure evidence
Enter fullscreen mode Exit fullscreen mode

The important part is not that code gets generated.

Plenty of systems can do that.

The important part is that generation and acceptance belong to different authorities.


The generator is not allowed to say "I passed"

This is the central rule.

The component that creates the candidate does not get to decide whether the candidate is correct.

Here is the core idea in one diagram:

Derivative / Forge separated authority: requirement, structured contract, candidate generation, sandbox execution, independent validation, and final packaging outcomes.

The important boundary is that packaging authority comes from validation evidence, not from the code generator itself.

The planner cannot declare correctness.

The coder cannot declare correctness.

Only the validation evidence can authorize packaging.

Internally, the software-building pipeline is called Forge.

Forge is built on top of the broader Derivative reasoning substrate.

Their responsibilities are deliberately separated:

Layer Responsibility
Derivative Constraints, deterministic reasoning, obligations, execution grounding, contradiction witnesses, audit and memory
Forge Software-build contracts, candidate generation, isolated execution, validation, bounded repair and packaging

This separation matters because otherwise the system has an obvious conflict of interest.

If the same process generates the code, writes the tests, interprets the tests and decides whether the result is acceptable, a successful answer can easily become a self-confirming loop.

Derivative tries to break that loop.


Natural language becomes obligations

The first transformation is not:

prompt → code
Enter fullscreen mode Exit fullscreen mode

It is closer to:

requirement → explicit obligations → code
Enter fullscreen mode Exit fullscreen mode

The requirement compiler preserves individual pieces of user intent and turns them into structured constraints.

That can include things like:

  • required public interfaces;
  • required files or commands;
  • behavioral acceptance criteria;
  • persistence requirements;
  • security expectations;
  • observability requirements;
  • test obligations;
  • forbidden behavior.

The distinction is important.

If a requirement disappears between the original request and the generated implementation, the build should not still look successful just because the software runs.

The contract is frozen before validation.

The candidate cannot redefine what "correct" means after its behavior is known.


Execution happens outside the generator

Generated software is untrusted software.

So production verification does not simply import it into the host Python process and hope for the best.

Forge executes candidates inside an ephemeral Docker sandbox with:

  • no network access;
  • a read-only root filesystem;
  • resource limits;
  • controlled environment variables;
  • no inheritance of host credentials.

This gives validation a real execution boundary.

The system is not asking the model:

Does this code look like it should work?

It is asking the environment:

What actually happened when this artifact ran?

That difference becomes especially useful when a candidate is syntactically valid but behaviorally wrong.


Validation has multiple layers

A build does not become verified because one test returned zero.

Forge separates validation into different kinds of evidence.

At a high level:

candidate
   ↓
syntax / import / execution
   ↓
requirement and acceptance checks
   ↓
adversarial validation
   ↓
packaging decision
Enter fullscreen mode Exit fullscreen mode

All required layers must pass before packaging is authorized.

A candidate can therefore execute correctly and still fail the build.

That is intentional.

The current build outcomes are:

Outcome Meaning
verified Required execution, contract and adversarial gates passed
validation_failed A candidate exists, but the evidence does not justify packaging
infeasible_proven The original constraints are contradictory and the system produced an evidence-backed certificate

Operational failures are kept separate.

For example:

sandbox_unavailable
sandbox_policy_violation
Enter fullscreen mode Exit fullscreen mode

Those are not silently converted into failed software requirements.

They mean the evaluation itself could not legitimately happen.


verified does not mean "mathematically correct forever"

This is another boundary I wanted the project to make explicit.

In Derivative, verified does not mean:

  • formally proven;
  • universally correct;
  • secure against every possible attack;
  • correct for every future input;
  • equivalent to human review.

It means something narrower:

At this revision, the artifact satisfied the executable contracts and evidence checks that Forge knew how to apply.

That may sound less impressive than saying "verified software."

I think it is more useful.

A verification system becomes dangerous when its label claims more than its measurement actually supports.


Repair is allowed, but it is bounded

When validation finds a concrete failure, Forge can attempt a repair.

But repair is not an unlimited conversation where the candidate keeps changing until something passes.

Retries are tied to observed failure signatures.

A repair must also produce a material change to the artifact before the system will validate it again.

That keeps the loop closer to:

failure evidence
      ↓
targeted modification
      ↓
new artifact
      ↓
full revalidation
Enter fullscreen mode Exit fullscreen mode

rather than:

something failed
      ↓
keep trying random changes
      ↓
eventually declare success
Enter fullscreen mode Exit fullscreen mode

The evidence is part of the artifact

Every run produces structured evidence.

A successful build can contain artifacts such as:

build_spec.json
feasible_plan.json
code_artifact.json
validation_artifact.json
packaged_artifact.json
Enter fullscreen mode Exit fullscreen mode

A packaged result also retains information about the code, tests and validation that authorized its creation.

This is useful for two reasons.

First, the decision becomes inspectable.

Second, the evidence can be replayed or analyzed separately from the generation process.

The build is not just:

here are some files
Enter fullscreen mode Exit fullscreen mode

It is closer to:

here are the files
+
here is the contract they were evaluated against
+
here is what was executed
+
here is why packaging was allowed
Enter fullscreen mode Exit fullscreen mode

The uncomfortable part: the blind benchmark is not flattering

This is where the project became more interesting to me.

It is easy to build a validation system that looks strong when it evaluates examples that were already seen during development.

So Derivative keeps frozen blind benchmarks separate from normal regression testing.

The current frozen V11 baseline contains 12 cases.

The result was:

status accuracy:       6 / 12
external Verified@1:   0 / 6
Enter fullscreen mode Exit fullscreen mode

None of the six cases expected to produce verified software reached the external oracle.

At the same time:

validation_failed cases: 3 / 3 correct
infeasible cases:        3 / 3 correct
Enter fullscreen mode Exit fullscreen mode

That is obviously not a production-quality code-generation result.

But it revealed something important.

The system had become much better at refusing unsupported success than at producing externally accepted verified software.

For this project, that is useful information.

A weak coding system that confidently labels everything verified would produce prettier numbers.

It would also defeat the entire reason Derivative exists.

Frozen blind results remain frozen.

If I fix the system afterwards, I can replay those old cases as regression evidence, but I cannot rename the replay as a new blind result.

The next real measurement requires a new unseen distribution.


Failure can be a valid result

This idea appears throughout the project.

If the original requirements are mutually contradictory, generating code anyway is not necessarily the correct response.

Derivative can instead terminate with:

infeasible_proven
Enter fullscreen mode Exit fullscreen mode

and return evidence explaining the contradiction.

Likewise, if the software exists but the validation evidence is not strong enough:

validation_failed
Enter fullscreen mode Exit fullscreen mode

is a valid terminal state.

The goal is not to maximize how often the pipeline says yes.

The goal is to make the meaning of yes stronger.


This is not another general-purpose coding agent

The current scope is intentionally narrow.

Forge currently focuses on greenfield Python artifacts such as:

  • CLI applications;
  • REST services;
  • data pipelines;
  • libraries.

It does not currently claim:

  • arbitrary existing-repository editing;
  • frontend generation;
  • multiple programming languages;
  • universal formal verification;
  • general software correctness.

That limitation is deliberate.

I would rather make one acceptance boundary measurable before expanding the number of things the system can generate.


Why I built it this way

Modern coding models can generate increasingly large amounts of plausible software.

That changes the bottleneck.

Producing code is becoming cheaper.

Determining what deserves to be trusted is not.

When generation becomes abundant, the interesting engineering problem shifts toward:

  • provenance;
  • execution;
  • independent checks;
  • explicit assumptions;
  • failure semantics;
  • acceptance authority.

That is the part I wanted to experiment with.

Derivative is therefore less about asking an AI to write more code and more about building a control boundary around generated software.

The generator proposes.

The runtime produces evidence.

The validator decides whether the evidence is sufficient.

And sometimes the correct output is simply:

no
Enter fullscreen mode Exit fullscreen mode

Where the project goes next

The current phase is deliberately focused on the verification mechanism rather than expanding into more languages or domains.

The next useful progress is not another feature list.

It is improving the distance between:

internally verified
Enter fullscreen mode Exit fullscreen mode

and:

accepted by an independent external oracle
Enter fullscreen mode Exit fullscreen mode

without weakening the conditions required for verified.

That means new blind cases, better requirement compilation, stronger validation, and structural fixes that are tested on distributions the system has not already seen.

The important constraint remains the same:

a generated artifact does not get to certify itself.

That rule is simple enough to explain in one sentence.

Making it work reliably turned out to be a much larger software problem.


Derivative is open source:

github.com/Daniele-Cangi/Derivative

Top comments (0)