DEV Community

Cover image for The Contract Discovery Bottleneck

The Contract Discovery Bottleneck

Ken W Alger on September 10, 2026

AI can generate the code. We can verify the behavior. But who decides what correct means? I wrote recently about a coding agent that built me a pa...
Collapse
 
alexshev profile image
Alex Shev •

This makes the specification problem concrete: the behavior people call obvious is usually absent from the artifact that code and tests actually consume. I like the idea of treating those negative cases as first-class requirements, because they give reviewers something precise to challenge before implementation hardens around an assumption.

Collapse
 
joinwell52 profile image
joinwell52 •

Glad to see the concurrent-reset case in this follow-up. One detail I’d want to keep is the contract version attached to each verification result. After adding that concurrency rule, an older green run would still tell us something useful—just not that the stronger contract had been checked.

Collapse
 
kenwalger profile image
Ken W Alger •

Yes. I think that's an important consequence of calling the specification durable and allowing it to evolve.

An old green result shouldn't become false when the contract changes. It remains evidence that implementation X satisfied contract version Y at that point in time. It cannot establish compliance with the stronger version Y+1.

So the verification result probably needs to bind together at least the implementation, contract version, result, and time of verification. Otherwise "passed" loses the context necessary to know what was actually proven.

And now we've wandered rather directly into provenance. :)

Collapse
 
vinhnguyenthanhdn profile image
Vinh Nguyen •

The reset case may be less about the specification being incomplete than about the shape the invariant was written in. "A second attempt is rejected" is a per-request predicate, and I ran the race to see what it actually binds to: across 200 trials of a non-atomic check-then-set, all 200 produced two successful resets, and in not one of them was any request rejected. So under concurrency that sequential assertion does not fail, it has nothing to assert on, because there is no second attempt that ever observes a used token. Write the same sentence as a count over the token's lifetime instead, successes <= 1, and it fails on the first trial, while the atomic version passed 200 of 200 with exactly one rejection each. That bears on the property-based question you close with, since a count-shaped invariant does not require anyone to have thought of concurrency first, whereas the predicate form is quietly satisfiable by both branches of the race.

Collapse
 
kenwalger profile image
Ken W Alger •

That's a really useful distinction, and I think you're right that it puts pressure on how I framed the missing invariant.

successes <= 1 describes the property I actually care about across the token's lifetime without requiring the specification author to predict concurrency as the failure mechanism. "A second attempt is rejected" accidentally embeds a sequential model into the requirement.

That also makes the property-based testing question more interesting than I gave it credit for. The framework still can't invent an unstated requirement, but the shape of a well-stated invariant can expose failure modes the author never anticipated.

And 200/200 is a fairly persuasive way to make the point. :) Thanks for actually running it.

Collapse
 
aiops-enabler profile image
AiOps Enabler •

The line I keep rereading is where you say the system needs a definition of what the agent is permitted to cause and what evidence would establish that the intended effect actually happened, and that you don't have the architecture for it yet. I don't either, but I have two production scars that mark where the hole is.

For generated code the evidence channel is free: the artifact is sitting there and you can rerun it. For an action, the record IS the artifact. That makes the emitting path load-bearing, and it fails in two directions that both look fine.

Evidence without execution. Our onboarding wizard generated a reporter that posted outcome: success on a 30-minute timer whether or not the agent had run. Weeks of clean, continuous, well-formed, entirely fictional evidence. Every contract you could write over that data was satisfied. Nobody flagged it, because flagging requires a reason to look and a green record gives you none.

Execution without evidence. Our run attestations were bound to (repository, workflow filename). Someone renamed a workflow, and reporting began 404ing after the work had already been done. Authorised, executing, invisible. A quiet agent is indistinguishable from an idle one.

Collapse
 
kenwalger profile image
Ken W Alger •

Those are fantastic counterexamples because they fail in opposite directions, even though both produce apparently reasonable system states.

The first one especially breaks a naive version of what I wrote. If evidence can be emitted independently of execution, then no contract over the evidence can prove the action occurred. You can have perfectly valid, well-formed, cryptographically pristine evidence of something that never happened.

And the workflow rename gives you the inverse: execution occurred, but the evidence path was detached from it, so the action becomes invisible.

That makes me think the missing architectural requirement is stronger than "produce evidence after an action." Evidence generation has to be causally bound to the execution path somehow. The system shouldn't be able to produce an execution receipt without executing, and ideally shouldn't be able to execute without producing the corresponding receipt.

Which immediately raises harder questions about partial failure, of course. What happens if the action commits and evidence emission fails? At that point, the evidence path itself starts looking like part of the transaction rather than downstream observability.

These are exactly the kinds of production scars I hoped somebody would bring up in the part where I admitted I didn't have the architecture yet. Thanks for sharing them.

Collapse
 
kevinbai profile image
kevinbai •

The concurrency case has a mechanical fix that also sharpens the spec: make consumption atomic in the store instead of in the application protocol. UPDATE reset_tokens SET consumed_at = now() WHERE token = $1 AND consumed_at IS NULL, then reject if affected_rows is 0. "Single use" stops being a rule the app has to remember and becomes an invariant the database enforces on every competing request at once — two simultaneous attempts can't both succeed, because the second UPDATE matches zero rows. The spec sentence and the mechanism collapse into one statement.

The harder problem in your article is the shared answer key. Independence has to come from the definition of correct behavior being written down before the implementation exists. Otherwise the verifier and the implementation are two students who studied from the same incomplete notes, and their agreement carries no information.

Collapse
 
kenwalger profile image
Ken W Alger •

Yes, push enforcement down to the store. An atomic conditional update is a much stronger mechanism than asking application code to remember the rule correctly across competing requests.

I'd keep one small separation, though: I don't think the specification and mechanism quite collapse into the same statement. successes <= 1 is the behavioral invariant; the conditional UPDATE is one implementation that can enforce it. If I replace the database tomorrow, I still want the invariant to survive even though that mechanism disappears.

And I agree that the shared answer key remains the harder problem. The database can enforce single use beautifully once we've decided single use belongs in the definition of correct. It can't tell us that we forgot to require it.

Collapse
 
thegm26 profile image
George Michalakis • • Edited

Honestly, I feel like LLMs promised us a less painful future and I find myself entangled into a higher abstraction layer of how to know the ins and outs functionally and more importantly error-prone wise of huge chunks of "foreign" code.

Ughhh

Collapse
 
kenwalger profile image
Ken W Alger •

I feel some of that tension too. :) We were promised less time wrestling with implementation, and in some ways we've gotten it. But the work hasn't necessarily disappeared. Some of it has moved upward into understanding behavior, boundaries, failure modes, and code we didn't personally write.

What worries me is when generation speed outruns our ability to build a mental model of what was generated. At that point we're not really maintaining code we understand so much as auditing an increasingly large foreign artifact.

Maybe the win isn't eliminating that higher-level reasoning, but making sure we're spending human attention on judgment and invariants rather than boilerplate. I'm not convinced we've figured out that balance yet either.

Collapse
 
jo-do profile image
Jo Do •

The independence of the spec is doing all the work here. If the spec gets written from the same conversation that produced the prompt, spec and implementation share the same blind spot - you end up with two artifacts that agree with each other and are both wrong, which is worse than one. The single-use reset link is the perfect example because "single use" lives in security policy, not in the feature request. Nobody asks for it because it feels obvious, and it stays obvious right up until the day it isn't. We started writing acceptance criteria from the ticket before touching the prompt at all, specifically so the spec can't absorb the implementation's assumptions. The closer the spec writer's information source is to the generator's, the less the verification step buys you.

Collapse
 
kenwalger profile image
Ken W Alger •

Yes, and I think your last sentence gets closer to the real independence requirement than simply saying “use a different agent” or “write the tests separately.”

Two artifacts can be produced independently and still inherit the same blind spot if their information lineage is effectively identical. In that case, we've separated the work without introducing any new knowledge.

Writing acceptance criteria from the ticket before constructing the implementation prompt is interesting because it deliberately creates some distance between those paths. I'd still want other sources feeding the contract too, especially security policy, known domain invariants, previous incidents, and the accumulated “obvious” constraints that rarely make it into feature requests.

I'm increasingly thinking independence needs to be evaluated by where the knowledge came from, not simply who or what produced the artifact. Thanks for sharpening that distinction.