1. Introduction
When I first started building my AI-powered course recommendation system, I thought integrating an LLM with backend APIs would be straightforward.
However, I quickly ran into a key problem:
The model was calling backend APIs almost every time—even when it wasn’t necessary.
This led to:
- Increased latency
- Unnecessary API calls
- Inefficient system behavior
In this post, I’ll walk through how I redesigned my AI Agent to behave more intelligently using controlled tool-calling and multi-step reasoning.
2. The Problem
2.1 Initial Design (Naive Approach)
My initial architecture looked like this:
User Input → LLM → Function Call → Backend API → LLM → Response
The idea was simple:
- Let the model decide which function to call
- Always execute the function if suggested
2.2 What Went Wrong
In practice, this caused several issues:
- The model triggered function calls even for simple messages like “Hello”
- Redundant API calls increased backend load
- No clear control over when tools should be executed
- Poor user experience due to unnecessary delays
At this point, I realized:
Letting the LLM fully control execution without constraints leads to inefficient systems.
3. Key Insight
The turning point was understanding this:
An AI Agent should not just “call tools” — it should decide when NOT to call them.
This meant I needed:
- A decision layer
- Controlled execution logic
- Better orchestration between LLM and backend
4. System Redesign
4.1 New Architecture
Instead of blindly executing tool calls, I redesigned the system:
User Input
→ LLM (intent + decision)
→ Orchestration Layer
→ (Conditional) Backend Tool Execution
→ LLM Final Response
4.2 Decision-Based Tool Calling
I introduced logic where:
- Simple inputs → direct response (no tool call)
- Informational queries → fetch data via API
- Complex actions → multi-step reasoning before execution
Example:
| Input | Behavior |
|---|---|
| “Hello” | No tool call |
| “What courses do you have?” | Call course API |
| “Enroll me in a backend course” | Gather info → then call tool |
4.3 Deferred Execution (Important)
Instead of calling tools immediately, the agent:
- Identifies missing information
- Asks follow-up questions
- Executes the function only when all parameters are available
This significantly reduced unnecessary calls.
5. Backend as Tools
On the backend (ASP.NET Core), I designed APIs as callable tools:
- GetCourses()
- ValidateUser()
- EnrollCourse()
Each function was exposed with structured input/output schemas so the LLM could interact with them safely.
6. Prompt & Orchestration Strategy
To improve behavior, I refined:
- System prompts (clear tool usage rules)
- Function descriptions (explicit intent)
- Context handling (multi-turn conversations)
Goal:
- Improve intent recognition
- Reduce incorrect tool selection
- Maintain consistent responses
7. Evaluation & Testing
Since I didn’t have large-scale production traffic, I evaluated the system using:
- Simulated user inputs (various intent scenarios)
- Multi-turn conversation testing
- Edge cases (incomplete or ambiguous queries)
Key observations:
- Reduced unnecessary tool calls
- Improved response consistency
- Better handling of complex requests
8. Trade-offs
Every design has trade-offs:
Pros:
- More efficient API usage
- Better user experience
- Clearer control over system behavior
Cons:
- Increased system complexity
- Requires careful prompt and flow design
- Harder to debug than simple chatbot systems
9. What I Learned
This project changed how I think about AI systems:
- LLMs should be controllers, not executors
- Backend systems should be tools, not just APIs
- Good AI systems require orchestration, not just prompts
10. Conclusion
Building an AI Agent is not just about calling an API.
It’s about designing a system where:
- Decisions are controlled
- Execution is efficient
- Behavior is predictable
If you’re building AI-powered applications, focus less on “what the model can do”
and more on how your system controls it.
11. Future Improvements
- Add conversation memory (persistent context)
- Integrate vector search for better recommendations
- Introduce performance metrics (latency, tool-call rate)
- Optimize system for real-world scaling
12. Final Thoughts
This project pushed me beyond just using LLMs —
it helped me think like an engineer designing systems around them.
And that’s where real value comes from.
Top comments (1)
The deferred execution pattern for mutating calls like enrollment saves a lot of headache. When developers give agents write access to backend endpoints, the model almost always attempts to call the function with placeholder or missing arguments on the very first turn.
One operational adjustment that complements that decision layer is dynamic tool masking in the request payload itself. If you include every endpoint schema in every turn, you pay the token overhead for the full tool manifest on simple greetings, and the model has a non-zero chance of reaching for them regardless of system prompt warnings. Suppressing tool definitions entirely or toggling tool_choice to none on the initial classification pass drops both the round-trip latency and the prompt token cost before the orchestration layer even runs.