DEV Community

Cover image for Why AI Coding Agents Still Get APIs Wrong Even When They Have Documentation
Saurabh Yadav
Saurabh Yadav

Posted on

Why AI Coding Agents Still Get APIs Wrong Even When They Have Documentation

AI coding agents have gotten very good at writing code.

But APIs are still a surprisingly easy place for them to make mistakes.

The problem isn't always that the model doesn't know the API.

Sometimes the documentation is:

  • for a different version
  • spread across multiple pages
  • buried inside a traditional documentation site
  • missing from the model's context
  • inconsistent with examples found elsewhere
  • full of deprecated endpoints

That's what led me to build DocOrbit.

DocOrbit is an open-source documentation intelligence layer for coding agents.

Instead of treating documentation as a giant block of text, the goal is to turn it into structured, version-aware evidence.

The pipeline looks roughly like:

Documentation discovery -> Documentation graph -> Project dependency detection -> Version resolution -> Task-specific retrieval -> API intelligence -> Implementation context -> Code/API verification

One part I'm particularly interested in is verification.

If an agent generates code using the wrong endpoint, wrong HTTP method, missing required parameters, or a deprecated API, DocOrbit can compare the implementation against the documented contract and report the result.

The project is still evolving, but the goal is simple:

Don't just give coding agents documentation. Give them the right documentation, and check whether their implementation agrees with it.

GitHub: https://github.com/HakashiKatake/docorbit

I'd love feedback from people building coding agents, MCP servers, developer tools, or documentation infrastructure.

Top comments (3)

Collapse
 
tercelyi profile image
tercel •

The “wrong endpoint / method / params / deprecated API” list is the painful core here, because it spells out what’s really happening: models are guessing against a fuzzy mental mashup of docs instead of being checked against a crisp contract.

Your pipeline step “Documentation graph → Project dependency detection → Version resolution → Task-specific retrieval” basically implies: the agent should be grounded in one specific, versioned truth before it ever types client.foo().

A few questions I’d be curious about:

  • How strict is the “API intelligence” layer? Does it normalize things into an IR like OpenAPI/JSON Schema, or does it work directly on semi-structured docs?
  • In verification, do you only do static checks (signatures, required params), or do you also try “plausible” runtime checks, like building sample calls from docs and comparing?
  • How do you handle gray areas, like optional-but-actually-required params that the docs don’t admit? Does DocOrbit report “contract says OK, but examples disagree”?

The version-resolution piece feels huge for MCP-style tools in particular, where agents might talk to multiple tools each with its own evolving surface.

I’d love to see examples of DocOrbit’s failure reports: what does the agent get back, and how does that feed into self-correction vs. just surfacing an error to the user?

Collapse
 
jo-do profile image
Jo Do •

The version-mismatch failure is the one that burns the most time because the agent looks confident and the code looks right - it is right, for a version you do not run. "Documentation exists" and "the model is grounded on the documentation that matters" are completely different states, and scattered docs make the second one accidental. The fix pattern that works for me is pinning the exact spec version into context rather than pointing at a docs site: the model will happily average five versions of an API into one fluent, wrong call.

Collapse
 
mervyx profile image
MERVYX •

Version-aware evidence is the key distinction here. One practical addition I’ve found useful is to make the verification output explicit: requested version, resolved source URL/commit, endpoint/method/required params, and status (VERIFIED / NEEDS_CHECK / NOT_FOUND). That turns “docs were available” into an artifact you can audit without claiming the whole integration is correct. MERVYX is exploring a market where people and AI/Agents buy and sell digital capabilities, tasks, and outcomes; we're careful to keep proposed checks separate from implemented services. In DocOrbit, which verification artifact is hardest to keep accurate as docs drift—version resolution, dependency detection, or the final code/API comparison?