DEV Community

Jeff
Jeff

Posted on Originally published at powerduck.com

AST vs AI for reverse-engineering OpenAPI: where static analysis stops and the model earns its place

Two pitches dominate the "generate API docs from code" space. One says "point an LLM at the repo, it will understand everything." The other says "static analysis is deterministic, just parse the files." Both fail on real backends, for opposite reasons. The designs that actually hold up in production use a strict division of labor: static analysis owns the closed world of routes and types, and the model is allowed to touch only the narrow pieces that static analysis can prove it cannot know.

Why the all-AI approach fails

Feeding a whole repository to a chat model and asking for OpenAPI is impressive on a three-route demo and dangerous on a real service.

  • It invents. Models surface plausible endpoints and fields that do not exist, and omit unglamorous ones that do. There is no mechanism tying an emitted operation to evidence in the code.
  • It is not reproducible. The same code scanned twice can yield different required fields and enums, which makes the output impossible to gate in CI.
  • It over-exposes. Asked for a response shape, a model tends to serialize the entire database entity, leaking columns the endpoint never returns, including internal flags and credential-adjacent fields.
  • It leaks source. The entire codebase, not the relevant contract, is sent to a third party, which is a non-starter for proprietary backends.
  • It hides uncertainty. A confident-looking schema gives no signal about which fields were proven and which were guessed.

An OpenAPI document is a contract. Contracts need evidence and repeatability, neither of which is a free-form model's strength.

Why pure static analysis goes blind

The opposite extreme, regexes and AST walks with no semantic layer, fails the other way. Dynamic frameworks construct responses in helpers, serializers, and mappers that a syntax-only pass cannot follow:

  • A computed serializer that assembles a response from several services has a return shape the parser cannot fully resolve.
  • A Django SerializerMethodField or a runtime-built choices list has no static enum.
  • A response constructed across a cross-package helper, with error and success branches, leaves the success schema partially unknown.
  • Query parameters coerced through hand-rolled parsing accept a wider set than a naive type guess suggests.

A tool that refuses to emit anything it cannot fully prove produces documents full of unknown, and teams stop using them. A tool that fills those holes with assumptions produces documents that lie. The honest answer is a third state, plus a constrained way to resolve it.

The pipeline that works

Reliable scanners separate the deterministic core from the optional intelligence:

indexer
  -> language packs   (AST plus the type checker where one exists)
  -> framework packs  (route graphs, handlers, instance tracing)
  -> completeness gate (proven / proven-absent / gap)
  -> Discovery IR
  -> validated OpenAPI 3.2
Enter fullscreen mode Exit fullscreen mode

Routes are a closed world established by tracing the real framework instance, so lookalike calls such as cache.get(...) are never mistaken for endpoints and unmounted routers are reported as unreachable. TypeScript projects use the compiler checker to resolve generics, named interfaces, enums, utility types, and validation schemas; the other languages are parsed through tree-sitter WASM with no language toolchain installed. After that pass, every parameter, request body, and response is in exactly one of three states: proven with evidence, proven absent, or an explicit gap carrying a code such as query-unknown, body-schema-unknown, response-unknown, auth-unknown, or sse-events-unknown.

Only the third state is eligible for AI.

Three rules that keep the model honest

When a gap resolver is enabled, three constraints prevent it from becoming an all-AI generator in disguise.

1. It can never invent a route, method, or path. The set of operations is fixed by framework tracing before any model call. The resolver fills properties of an operation the scanner already proved exists. It cannot add an endpoint no matter how likely one seems.

2. It sees the smallest slice that contains the gap, never whole files. For a handler with an unproven response schema, the model receives that handler's slice and the specific gap code, not the surrounding module, other routes, or unrelated secrets. The scanning package itself never calls a model vendor; the host application makes the call, behind explicit opt-in, using the user's own model configuration. Source code is not sent anywhere by default.

3. Its output is clamped and reviewable. The response is validated against a strict JSON Schema subset: no $refs, bounded depth and property counts. The resolver can fill query parameters, headers, request bodies, status-keyed response schemas, and SSE event payloads. In the desktop workflow each proposal is surfaced as a review item that a person accepts, edits, or rejects. A failed or rejected fill is never fatal and never hidden; the gap stays visible in the report.

A failed fill returning null is a good outcome. It means an unknown contract stayed labeled unknown instead of being promoted to a confident guess.

Reproducibility is a feature, not an accident

For AI-assisted scanning to belong in CI, the same inputs must produce the same result. Deterministic extraction already does. The model step is pinned with a prompt-contract version and a cache keyed on the handler slice and prompt version, and it is called at temperature zero, so an unchanged handler reuses its prior resolution. Provenance lives in a .powerduck/discovery.json sidecar rather than in the generated document, which keeps the OpenAPI itself clean and editable. On rescan, a three-way merge applies fresh evidence while treating manual edits as authoritative: changed routes refresh structural contracts but preserve descriptions, examples, tags, and extensions, and removed routes are flagged rather than silently deleted. The human always owns the document.

Where the model genuinely earns its place

The model is valuable precisely where static proof is impossible without executing the code: inferring the output of a computed mapper, proposing a schema for a SerializerMethodField, or suggesting the shape behind a cross-package helper. Treat those proposals as hypotheses. Each one is either corroborated by a real request against a fixture, downgraded to an explicit unknown, or corrected by a human. The strongest pipeline runs all three layers in order: static proof first, constrained AI for the residual gaps, and live requests as the final arbiter when a contract cannot be proven from source.

Choosing your default

  • Use static-only output for CI gating and any repository where source cannot leave the machine; the gap report is still a precise worklist.
  • Enable AI gap fills when you want a first draft of the dynamic pieces and are prepared to review them; keep the review step mandatory.
  • Reach for runtime verification on anything security- or billing-adjacent, where a wrong field is more expensive than an honest unknown.

The mindset shift is to stop measuring a scanner by how complete its output looks and start measuring it by whether every claim is either backed by evidence or clearly labeled as unresolved. A document with ten visible gaps is more trustworthy than one with ten silently invented fields.

Powerduck implements this exact split: deterministic framework tracing and type analysis by default, with opt-in, reviewable AI gap fills that never create routes, plus incremental rescans that preserve your edits. See it end to end in the online demo, and compare the static extraction across languages in the code-to-OpenAPI overview.

Top comments (0)