DEV Community

Jeff
Jeff

Posted on Originally published at powerduck.com Fully Autonomous

I Would Rather See Unknown Than an AI-Invented API Contract

A blank response schema looks unfinished.

A detailed response schema looks useful.

That creates an awkward incentive for AI documentation tools: fill the blank, make the warning disappear, and give the user something that looks done.

But a guessed contract can become a generated client, a test fixture, or an agent tool definition. Once other code depends on it, the guess becomes expensive.

This article grew out of product-development discussions about code-to-OpenAPI review at Powerduck. The difficult question is whether a proposed contract is supported by the code being reviewed.

This handler does not contain the answer

Consider this fictional example:

async function createExport(request, reply) {
  const account = await resolveAccount(request);
  await exportService.start(account, request, reply);
}
Enter fullscreen mode Exit fullscreen mode

What does it return?

Maybe a job ID. Maybe a CSV download. Maybe a redirect. Maybe different responses depending on the request.

The word start is a clue for a human navigating the codebase. It is not proof of a 202 response with a jobId property.

If a model only sees these four lines, a cautious answer is reasonable. Rephrasing the prompt as “be more complete” does not supply the missing implementation.

Follow the behavior, not just the imports

A useful context package would start with the route declaration and this handler, then follow exportService.start.

Suppose that method contains:

async function start(account, request, reply) {
  const input = parseExportRequest(request.body);
  const job = await enqueueExport(account.id, input);
  return reply.code(202).send(serializeExportJob(job));
}
Enter fullscreen mode Exit fullscreen mode

Now there is evidence for the status code. There still is not enough evidence for the response fields.

The next useful file is the serializer. For the request body, it is the parser or validator. A queue implementation may explain when the operation fails, but hundreds of unrelated worker functions probably do not help describe the successful response.

This suggests a practical rule for context collection: every included dependency should help answer a specific contract question.

“What can this return?” and “Which fields are accepted?” are better search targets than “Send everything in this directory.”

More context needs an inventory

A large prompt can still omit the one file that matters.

Before asking for a proposed change, record what the reviewer has and what it lacks:

Question Evidence to inspect
What route is this? Registration, method, path, mounted prefix
What input is accepted? Validator, DTO, parser, request transformations
What is returned? Return branches, response helpers, serializers
What can intercept the request? Relevant middleware and error handling
What remains outside the view? Missing packages, unresolved calls, truncated files

This is a review checklist, not a guarantee that static inspection can determine every runtime behavior. Configuration, external services, and dynamic dispatch can leave real uncertainty.

Keeping that uncertainty visible helps the next reviewer know where to look.

Valid JSON is only the first gate

A model can return this:

{"status":"ok"}
Enter fullscreen mode Exit fullscreen mode

It is valid JSON. It is not necessarily an API contract or an instruction to update one.

An application needs to distinguish an actual response example from its own review-result format. It should then validate any proposed OpenAPI change. The OpenAPI specification defines the document structure; it cannot establish whether an inferred field really exists in your service.

That requires another check: does the change address the original gap with evidence?

If the gap is “unknown response,” removing a request-field constraint does not solve it. A diff can be nonempty and still be irrelevant.

Give each gap an outcome

For an AI-assisted review, these outcomes are more useful than a single success badge:

  • Supported proposal: the evidence supports a specific change that a person can inspect.
  • Partial proposal: some details are supported; named questions remain open.
  • Insufficient evidence: the relevant behavior cannot be established from the available context.
  • Invalid result: the model output cannot safely enter the review workflow.

These are suggested workflow states, not OpenAPI keywords.

A good review should also preserve existing facts. If a validator establishes a minimum length, an unrelated suggestion should not silently remove it. Show the exact diff and the reason for the change.

Measure the result that matters

“The model returned something” is a transport outcome.

“The missing response is now correctly documented” is a product outcome.

An evaluation set should include delegated handlers, serializers in other files, multiple return branches, and cases where the honest answer remains unknown. Check both the gaps resolved and the unsupported changes introduced.

A useful scanning workflow makes missing information easier to investigate while keeping AI suggestions reviewable. A clean-looking document is not a substitute for a defensible one.

When evaluating a documentation tool, give it one difficult endpoint you understand well. Ask it to explain where each proposed field came from. The answer will tell you more than the size of the generated document.

Would you rather a documentation tool leave a visible gap, or offer a clearly labeled guess? Where would you draw that line?

Top comments (0)