DEV Community

Jeff
Jeff

Posted on

Day 6: I connected my coding agent to the API spec over MCP

Watch a coding agent integrate with an API it has never seen and you can predict the failure with eerie accuracy. It builds the screen fast, wires the form, posts to the create endpoint, and the integration is wrong: limit became pageSize, the list unwrapped data.records instead of data.items, and the amount field is a float. The agent is not being stupid. It is making the statistically most likely guess about a contract it was never shown.

Day 6 of the the series: stop pasting the API into chat, and serve the spec to the agent as tools it can discover and call.

Pasted context is lossy and temporary

The standard ritual — copy the docs, paste the relevant schema, warn about the weird bits — fails in three predictable ways:

  1. It compresses. A three-thousand-line spec becomes a three-hundred-line summary, and the field that breaks your feature is always in the omitted ninety percent.
  2. It decays. The paste captures the contract on Monday. The contract changes on Wednesday. The agent still has Monday.
  3. It does not survive context. Switch models, start a new session, hand off to a teammate, and the ritual restarts from zero.

The Model Context Protocol fixes the transport half of this: instead of pushing context at the model, the model pulls the exact operation, schema, or example it needs, when it needs it, in structured form. The other half is having a contract worth serving — which is what Days 2 through 5 produced.

What serving a spec over MCP actually gives the agent

An OpenAPI-to-MCP conversion turns the document into MCP tools, reusing the operation parameters, request and response schemas, descriptions, and referenced components. When the agent implements GET /holds/{holdId}, it does not infer the identifier — it retrieves a tool definition that says the path takes a UUID holdId, the response is a Hold with a status enum of held | confirmed | released | expired, and the call requires a bearer token. It cannot paraphrase a field name away because it never types the field name; the schema does.

Three categories of content flow through, and all three matter:

  • Tools — the operations themselves, with parameters and responses the agent validates against.
  • Resources — the raw spec sections, so the agent can read the error envelope or the pagination convention in the document's own words.
  • Prompts — team-authored guidance such as "money is always integer cents" or "all writes require Idempotency-Key," attached to the contract instead of repeated in chat.

That last category is where institutional knowledge finally lives somewhere durable.

Widen the surface deliberately

If the API is internal and the agent has credentials, exposing writes is convenient. If the spec is being served to partners or to hosted agents, start read-only and treat write operations as a separate decision:

  1. Expose list and detail operations first.
  2. Verify one real query end to end: agent discovers the tool, calls it against the running service, and answers a question from the actual response.
  3. Add writes operation by operation, each with its own confirmation and permission design.

MCP solves discovery and transport. It does not replace the service's authorization — a tool being discoverable must never imply that every caller may execute it. Upstream auth, MCP access keys, and per-operation exposure are separate gates on purpose.

Generate code in the order that fails cheaply

"Implement the API" as a single prompt yields forty endpoints of unverified assumptions. Slice vertically and let the agent generate against the served contract in this order:

Order Generate Why this order
1 Types and DTOs from schemas Mechanical, high-volume, everything compiles against them
2 Boundary validation Bad input fails at the edge with the documented error envelope
3 One happy path end to end Settles routing, DI, transactions, serialization for every later route
4 Documented failure cases 409 on duplicate, 422 on bad input, idempotent replay
5 The remaining operations Copies of a proven slice instead of independent guesses

The vertical slice is the whole trick. One operation working against the real database and clock hands the agent a concrete pattern to imitate, which it does far better than it follows abstract instructions.

Verify, then feed failures back as evidence

This is where the loop from Day 4 pays off. Run the scenario suite against localhost. When generated code returns a bare string on 409 instead of the error envelope, do not describe the bug in prose — hand the agent the failed scenario: the request, the expected schema, the actual response. That is evidence, not opinion, and agents fix evidence reliably.

Then run the same suite in CI so a pull request that drifts from the contract fails the build. The common generated-code defects — renamed fields, missing status codes, happy-path-only handlers — are each one scenario assertion away from being caught.

The drift loop, closed

Code and spec diverge even under discipline; under AI velocity they diverge faster. The Day 5 scanner runs in the opposite direction on purpose. After a sprint of generated code, rescan, diff the discovered routes against the design spec, and reconcile:

  • a route in code but not in the spec is either an invented endpoint or a missed requirement;
  • a route in the spec but not in code is unfinished work, visible before release night;
  • a differing parameter is a broken contract, and the sidecar diff points at the file.

Design with AI, generate from the contract, verify with scenarios, rescan for drift. Crucially, none of the contract lives inside the assistant — it lives in the local file — so switching models or tools never means reintroducing the API. I made that argument in detail here, and the mechanics of turning existing endpoints into agent tools are walked through here. The workspace I use serves the spec locally for desktop agents and publishes a hosted MCP endpoint for partners; any OpenAPI-to-MCP converter in front of a disciplined spec gets you the same loop.

Tomorrow, the final day: getting the document — human docs and agent tools together — in front of partners without turning the release into a documentation website project.

Top comments (0)