DEV Community

Cover image for Computer-Use Agents and MCP: What Agent-to-Agent Commerce Reveals About Protocol Boundaries
mech.app
mech.app

Posted on Originally published at mech.app

Computer-Use Agents and MCP: What Agent-to-Agent Commerce Reveals About Protocol Boundaries

Computer-use agents are moving from demos to production, and the plumbing is more interesting than the headlines suggest. When an agent needs to call another agent or interact with external tools, the protocol layer matters. MCP (Model Context Protocol) is one answer, but it sits in a specific slice of the stack between the model's reasoning loop and the tool execution boundary.

The conversation between Chris Benson and Demetrios Brinkmann on Practical AI surfaces the architectural choices that matter when agents start talking to each other, especially around commerce, enterprise deployment, and failure recovery.

The Agent Harness Layer

An agent harness is the runtime wrapper that sits between the LLM and the outside world. It handles:

  • State management: Conversation history, tool call results, retry logic
  • Tool registration: Which functions the model can invoke and their schemas
  • Execution sandboxing: Where and how tool calls run (local process, container, remote API)
  • Observability hooks: Logging, tracing, and cost tracking per invocation

The harness is not the model. It's the orchestration layer that decides when to call the model, what context to pass, and how to handle the model's output (text, tool call, or error).

State Management Trade-offs

Approach State Location Failure Recovery Multi-Agent Coordination
In-memory Harness process Lost on crash Requires external sync
External store (Redis, Postgres) Centralized DB Survives restarts Shared state possible
Event log (Kafka, NATS) Append-only stream Replay from offset Natural audit trail
Hybrid (local + remote) Both Complex reconciliation Best observability

Most production harnesses use a hybrid model: ephemeral state in-memory for speed, durable checkpoints in Postgres or S3 for recovery.

MCP as the Agent-to-Agent Protocol

MCP defines how agents discover and invoke tools across process boundaries. It's not a replacement for HTTP or gRPC. It's a schema layer on top that standardizes:

  • Tool discovery: JSON schemas for available functions
  • Request/response format: Structured tool calls and results
  • Authentication handshake: OAuth2, API keys, or custom tokens
  • Error semantics: Retryable vs. terminal failures

When Agent A calls Agent B via MCP, the protocol handles:

  1. Discovery: Agent A queries Agent B's MCP server for available tools
  2. Invocation: Agent A sends a tool call with typed parameters
  3. Execution: Agent B's harness runs the tool and returns structured output
  4. Continuation: Agent A's model receives the result and decides next action

What MCP Adds Over Plain HTTP

MCP servers expose a /tools endpoint that returns JSON schemas. This lets agents introspect capabilities at runtime instead of hardcoding API contracts.

{
  "tools": [
    {
      "name": "get_inventory",
      "description": "Fetch current inventory levels",
      "parameters": {
        "type": "object",
        "properties": {
          "warehouse_id": {"type": "string"},
          "sku": {"type": "string"}
        },
        "required": ["warehouse_id"]
      }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

The calling agent can pass this schema directly to the LLM's function-calling interface. The model generates a tool call, the harness validates it against the schema, and the MCP client sends it to the remote server.

This is cleaner than parsing OpenAPI specs or writing custom adapters for every API.

Agent-to-Agent Commerce: The Protocol Stress Test

When agents start paying each other for services, the protocol boundaries get real. Commerce introduces:

  • Authentication: Which agent is calling, and do they have credit?
  • Rate limiting: Per-agent quotas to prevent runaway costs
  • Idempotency: Retry logic without double-charging
  • Audit trails: Who called what, when, and how much it cost

MCP doesn't solve these problems out of the box. You need to layer on:

  • API gateway: Rate limiting, auth, and billing hooks
  • Idempotency keys: Client-generated UUIDs for safe retries
  • Event log: Immutable record of all agent interactions
  • Settlement layer: Reconcile usage with payment (Stripe, internal ledger, blockchain)

Example Flow: Agent A Buys Data from Agent B

  1. Agent A's harness generates a tool call: get_market_data(symbol="AAPL")
  2. Harness attaches an idempotency key and auth token
  3. MCP client sends request to Agent B's MCP server
  4. Agent B's gateway checks auth, rate limit, and balance
  5. If approved, Agent B's harness executes the tool
  6. Result returns to Agent A with a usage record
  7. Background job posts the charge to Agent A's account

If step 3 times out, Agent A's harness retries with the same idempotency key. Agent B's gateway sees the duplicate and returns the cached result without re-executing.

Where the Protocol Breaks Down

MCP works well for synchronous, short-lived tool calls. It struggles with:

  • Long-running jobs: If a tool takes 10 minutes, the HTTP connection times out
  • Streaming results: MCP doesn't define a streaming protocol (yet)
  • Stateful sessions: Each tool call is independent; no session affinity
  • Cross-agent transactions: No two-phase commit or distributed rollback

For long-running work, you need an async pattern:

  1. Agent A calls start_job(params) and gets a job ID
  2. Agent A polls get_job_status(job_id) until complete
  3. Agent A calls get_job_result(job_id) to fetch output

This is clunky but works. Some teams use webhooks: Agent B calls back to Agent A when the job finishes. That requires Agent A to expose an MCP server of its own, creating a bidirectional dependency.

Observability Across Agent Chains

When Agent A calls Agent B, which calls Agent C, you need distributed tracing. Each harness should:

  • Generate a trace ID on the first call
  • Propagate it in MCP request headers
  • Log all tool calls, results, and errors with the trace ID
  • Export traces to a collector (Jaeger, Honeycomb, Datadog)

Without this, debugging a failed agent chain is impossible. You see the final error but not which intermediate call failed or why.

Minimal Tracing Setup

import uuid
from opentelemetry import trace

tracer = trace.get_tracer(__name__)

def call_mcp_tool(server_url, tool_name, params):
    trace_id = str(uuid.uuid4())
    with tracer.start_as_current_span("mcp_call") as span:
        span.set_attribute("mcp.server", server_url)
        span.set_attribute("mcp.tool", tool_name)
        span.set_attribute("trace.id", trace_id)

        headers = {"X-Trace-ID": trace_id}
        response = requests.post(
            f"{server_url}/call",
            json={"tool": tool_name, "params": params},
            headers=headers
        )

        span.set_attribute("http.status", response.status_code)
        return response.json()
Enter fullscreen mode Exit fullscreen mode

Each MCP server should extract X-Trace-ID from incoming requests and include it in its own downstream calls.

Enterprise Deployment Challenges

Bringing computer-use agents into enterprise environments surfaces:

  • Security boundaries: Agents need scoped credentials, not root access
  • Compliance logging: Every action must be auditable for SOC2, GDPR, etc.
  • Failure isolation: One agent's crash shouldn't cascade
  • Cost control: Runaway loops can burn thousands in API calls

The harness is where you enforce these policies. Before executing a tool call, the harness checks:

  1. Does this agent have permission to call this tool?
  2. Is the call within rate limits and budget?
  3. Is the input sanitized (no SQL injection, path traversal, etc.)?
  4. Will the result be logged for audit?

If any check fails, the harness returns an error to the model instead of executing the tool.

The Harness-Model Relationship

The model generates tool calls. The harness decides whether to execute them. This separation is critical for safety.

Some teams use a "human-in-the-loop" harness that pauses before executing high-risk tools (delete database, send email, transfer money). The harness sends a notification to Slack or PagerDuty and waits for approval.

Other teams use a "guardrail" harness that runs a secondary model to validate tool calls. If the validator flags a call as risky, the harness rejects it and asks the primary model to try again.

Both patterns require the harness to maintain state across multiple model invocations. This is why in-memory state is fragile: if the harness crashes mid-approval, the context is lost.

Technical Verdict

Use MCP for agent-to-agent tool calls when:

  • You need runtime tool discovery instead of hardcoded API contracts
  • Your tools are synchronous and return in under 30 seconds
  • You want a standard schema layer that works with multiple LLM providers
  • You can layer on auth, rate limiting, and billing via an API gateway

Avoid MCP when:

  • You need streaming results or long-running jobs (use async patterns instead)
  • You require stateful sessions or distributed transactions (MCP is stateless)
  • Your agents run in high-security environments where dynamic tool discovery is a risk
  • You already have a mature API ecosystem and don't want to rewrite adapters

The agent harness is where the real work happens. MCP is just the wire protocol. Invest in harness observability, state management, and failure recovery before you scale agent-to-agent interactions.

Source Links

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to