DEV Community

Cover image for Most Developers Are Building AI Agents Wrong: MCP Is the Missing Contract
Sudhanshu Thakur
Sudhanshu Thakur

Posted on

Most Developers Are Building AI Agents Wrong: MCP Is the Missing Contract

AI agents are easy to demo and difficult to trust.

A developer can connect a language model to a few tools in an afternoon. The first demo looks impressive: the agent reads a request, calls an API, checks a database, and returns an answer.

Then production arrives.

The agent calls the wrong tool. It sends an incomplete payload. It retries a payment operation. It exposes information from another tenant. Nobody can explain why it made a decision.

The problem is not always the language model. The deeper problem is that many teams are building agents without a clear contract between the model, the tools, the data, and the people responsible for the outcome.

That is where the Model Context Protocol, or MCP, becomes important.

An AI agent is more than a chatbot

A useful enterprise agent usually has five parts:

  1. A model that reasons about the request
  2. Instructions that define its role and limits
  3. Tools that allow it to read or change systems
  4. Memory or context that helps it understand the task
  5. Guardrails that control what it is allowed to do

A chatbot mainly generates text. An agent can take actions.

That difference changes the engineering problem completely. If an agent only produces an explanation, an inaccurate response may be corrected by a human. If an agent can approve a refund, change a production configuration, or submit a healthcare request, an inaccurate response becomes an operational incident.

The agent must therefore be treated like a distributed system with an unreliable reasoning component.

The tool boundary is the real architecture

Many early agent implementations expose tools with a name, description, and input schema. That is not enough for production.

A safe tool contract must answer several questions:

  • Which user or tenant is the agent allowed to access?
  • Which values are valid?
  • Does the operation require approval?
  • Is the operation idempotent?
  • What happens if the network times out after success?
  • How is the action audited?

A tool description is not a security policy. It is only the beginning of a contract.

Why MCP matters

MCP provides a standard way for AI applications to discover and use external tools and resources. Instead of building a different integration for every model and every application, teams can expose capabilities through a consistent protocol.

That standardization is valuable, but it does not automatically make an integration safe.

MCP should be treated as an interoperability layer, not as permission to give an agent unrestricted access to a system.

A production MCP tool should define:

  • A clear business purpose
  • Read or write classification
  • Authentication and authorization requirements
  • Tenant and data-scope restrictions
  • Input validation
  • Idempotency behavior
  • Timeout and retry rules
  • Rate limits
  • Audit events
  • Human approval requirements
  • A safe failure response

The most important design principle is simple:

The model may request an action, but the platform must decide whether that action is allowed.

A safer agent flow

A reliable enterprise agent should follow a controlled sequence:

User request -> identity verification -> permitted context -> model plan -> policy evaluation -> human approval when required -> MCP tool execution -> result validation -> audit event -> user response

The model should not directly own the complete workflow. It should operate inside one.

For example, an agent may propose: Cancel this payment.

The policy layer should check whether the user is authenticated, whether the payment belongs to the user organization, whether it is still cancellable, whether the amount exceeds an approval threshold, whether the request was already processed, and whether the downstream tool is healthy.

Only after these checks should the action reach the payment service.

MCP and Spring Boot

A practical Java implementation can place an agent gateway in front of domain services.

The gateway can provide:

  • Spring Security authentication
  • OAuth2 or JWT validation
  • Tool-level authorization
  • Request and response schemas
  • Correlation IDs
  • Rate limiting
  • OpenTelemetry traces
  • Approval workflow integration
  • Central audit logging

The domain service should still enforce its own business rules. Never assume that because a request came from an AI gateway, it is trustworthy.

A safe execution model is: validate the user, evaluate policy, validate the tool input, enforce idempotency, call the domain service, validate the result, and record an audit event.

The model can select a tool. It should not bypass the same controls that apply to a human-operated API.

Read tools and write tools are not equal

Read-only tools may retrieve an invoice, search a knowledge base, or check an order status. Write tools may create a payment, update a customer, delete a document, or deploy software.

These categories should have different controls. A read tool may require normal user authorization. A high-impact write tool may require explicit confirmation, a second approver, a transaction limit, a dry-run mode, or a compensating action.

Do not hide a write operation behind a friendly name such as manage account. Use names that reveal the action and its risk: create refund, disable user, or deploy release.

Clear names help models, developers, reviewers, and security teams understand what is happening.

The biggest MCP security risks

MCP-based systems introduce familiar security risks in a new form.

Excessive permissions

An agent receives access to more data and tools than it needs. If the model or server is compromised, the blast radius becomes large.

Prompt injection

Untrusted content inside a document, webpage, or ticket attempts to influence the agent instructions. Retrieved content must be treated as data, not as a higher-priority instruction.

Confused deputy problems

The agent uses its own powerful credentials to perform an action on behalf of a user who does not have that permission.

Tool poisoning

A malicious or poorly maintained tool provides misleading descriptions, hidden side effects, or unsafe defaults.

Data leakage

The agent combines information from different users, tenants, or conversations and includes it in a response.

The solution is not one magic prompt. It is layered security: least privilege, isolation, validation, monitoring, and clear ownership.

Observability is not optional

Traditional application monitoring asks what request failed, which service was slow, and what status code was returned.

Agent observability needs additional questions:

  • What was the original user intent?
  • What context was retrieved?
  • Which tools were considered?
  • Which tool was selected?
  • What policy decision was made?
  • How many retries occurred?
  • Was human approval requested?
  • What data was returned to the user?

Every agent action should have a correlation ID that follows the request across the model gateway, MCP server, domain service, database, and audit system.

Without this trail, debugging becomes guesswork.

Start with narrow agents

The best first production agent is rarely a general-purpose assistant with access to the entire company.

Start with one bounded workflow:

  • Search internal documentation
  • Summarize support tickets
  • Validate an invoice
  • Classify incoming documents
  • Explain a failed transaction
  • Draft, but do not send, a customer response

Measure accuracy, latency, cost, refusal quality, tool-selection errors, and security violations. Then expand carefully.

A narrow agent with clear boundaries can create more business value than a general agent that appears intelligent but cannot be trusted.

The future of agents is controlled autonomy

AI agents will become more capable, but capability alone is not the goal.

The goal is controlled autonomy: systems that can complete useful work while remaining observable, explainable, reversible, and accountable.

MCP can help create a common interface between agents and the software world. But the protocol is only one layer. The real production architecture still needs identity, policy, security, resilience, testing, and human judgment.

The strongest agent teams will not ask only, What can the model do? They will also ask, What should the model be allowed to do, under which conditions, and how can we prove what happened?

That is the difference between an impressive demo and an enterprise system.

Final checklist

Before exposing an MCP tool to an AI agent, verify:

  • Is the tool purpose narrow and explicit?
  • Is access scoped to the authenticated user and tenant?
  • Are inputs validated independently of the model?
  • Are write actions idempotent?
  • Are high-risk actions reversible or approval-gated?
  • Are prompts and retrieved documents treated as untrusted data?
  • Are tool calls, policy decisions, and outputs audited?
  • Can the team trace a failed decision?
  • Is there a rate limit and emergency disable mechanism?
  • Does the system fail safely when the model is wrong?

If the answer to several of these questions is no, the agent is not ready for production.

It is ready for a better architecture.

A production agent architecture

flowchart LR
    U[User request] --> G[Agent gateway]
    G --> P[Policy and identity checks]
    P --> M[Model plan]
    M --> C[MCP server]
    C --> D[Domain service]
    D --> A[Audit and observability]
    A --> R[Validated response]

The model proposes the next action, but the gateway, policy layer, MCP server, and domain service each enforce their own boundary. This layered design prevents one incorrect model decision from becoming an uncontrolled production change.

Reference

Top comments (0)