DEV Community

Shaurya Singh
Shaurya Singh

Posted on

How We Stopped AI Agents from Hallucinating Tool Schemas & Wasting API Credits


If you’re building autonomous AI agents using frameworks like LangChain, CrewAI, AutoGen, or raw function calling, you’ve likely encountered this scenario:

Your agent executes Steps 1 through 4 flawlessly. Then, on Step 5, the model hallucinates a parameter type—for example, passing "amount": "$100" as a string instead of 100 as an integer.

The result?

Staging Data Corruption: The agent mutates production or staging environments with bad state.

Wasted API Credits: The agent hits rate limits and burns real money while stuck in broken execution loops.

Silent Runtime Failures: The target API returns a generic 400 Bad Request, giving the LLM zero actionable context on how to recover.

Here is an architectural walkthrough on how to fix this issue during local development using a virtual mock gateway.

The Root Cause: Syntactic vs. Semantic Schema Drift
When testing LLM tool calls against live backends, standard sampling temperatures (even low ones like 0.1) allow minor parameter drift.

Standard mock tools like Postman or WireMock expect deterministic HTTP requests. They fall short for non-deterministic AI agents because:

They don't validate incoming JSON arguments against strict AJV/Pydantic schemas on the fly.

They don't provide structured error feedback back to the LLM context to test self-correction.

The Architecture: Intercepting Tool Calls in Development
Instead of pointing your agent's @tool definitions directly to live external APIs (like Stripe, GitHub, or internal databases), route them through a virtual validation sandbox.

+------------------+ Malformed JSON Payload +-------------------+
| | ------------------------------------> | |
| AI Agent Loop | | MockAgent Gateway |
| (LangChain/Crew) | <------------------------------------ | (AJV Validation) |
+------------------+ 400 Bad Request + Error Path +-------------------+

  1. Define the Expected JSON Schema Set up a mock endpoint that expects strict data types. For instance:

JSON
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"properties": {
"user_id": { "type": "string" },
"amount": { "type": "integer", "minimum": 1 }
},
"required": ["user_id", "amount"]
}

  1. Live Payload Interception
    When the agent executes a tool call, the gateway checks the payload in real time before executing any code.

  2. Structured Self-Correction Feedback
    If the payload is invalid, do not return a generic string error. Return an explicit path error payload:

JSON
{
"status": "error",
"code": "SCHEMA_VIOLATION",
"details": [
{
"instancePath": "/amount",
"message": "must be integer",
"received": "$100"
}
],
"recovery_hint": "Cast the 'amount' parameter to a raw integer without currency symbols."
}
When the LLM reads this explicit path error in its execution history, its self-correction success rate on the following turn approaches ~90%.

Sandbox Solution: MockAgent
To manage this workflow without writing custom Express servers for every tool, check out MockAgent.

What it offers:

🛠️ Instant Mock Endpoints: Create virtual tool endpoints in seconds with custom JSON schemas.

⚡ Real-Time AJV Validation: Intercepts agent tool calls and checks arguments on the fly.

🔄 Self-Correction Support: Returns structured 400 errors so you can evaluate whether your prompt handles recovery gracefully.

📊 Trajectory Dashboard: Inspect real-time logs, execution timestamps, schema failure rates, and latencies.

Try the sandbox for free: https://mockagent.vercel.app/

Discussion
How are you currently handling schema drift and function-calling failures in your agent runtime? Are you enforcing Pydantic try/except retry blocks at the tool boundary, or relying purely on system prompt directives?

Let me know in the comments below!

Top comments (0)